Design method for generating multi-agent sensing integrated system
By combining generative multi-agent reinforcement learning algorithms and generative adversarial networks, the resource allocation challenge of multi-agent deep reinforcement learning in a cellular communication-integrated sensing system is solved, maximizing the sensing signal-to-interference-plus-noise ratio and reducing computational complexity, thereby improving system performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DALIAN UNIV OF TECH
- Filing Date
- 2026-01-30
- Publication Date
- 2026-04-28
AI Technical Summary
Existing multi-agent deep reinforcement learning faces challenges such as low sampling efficiency and unstable training in decellularized sensing integrated systems, especially in dynamic environments where it is difficult to effectively solve resource allocation problems.
A generative multi-agent reinforcement learning algorithm is adopted, which is combined with generative adversarial networks (GANs) to enhance multi-agent deep reinforcement learning (MADRL). Through centralized training and decentralized decision-making, the generative adversarial network is integrated into the multi-agent deep reinforcement learning. A joint optimization method for transmit beamforming, receive filtering and base station scheduling is designed and transformed into a Markov decision process to realize multi-station cooperation.
This improves the ability to maximize the signal-to-interference-plus-noise ratio (SINR), reduces computational complexity, enhances Q-value estimation accuracy and decision quality, and enables the design of a high-performance multi-station integrated communication and sensing system.
Smart Images

Figure CN121940783A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of multi-station collaboration in cellular communication and sensing integration. It designs a method for designing a generative multi-agent integrated sensing system. Specifically, it designs a low-complexity, high-efficiency collaboration scheme for multi-station communication and sensing integration based on generative multi-agent reinforcement learning. In particular, each communication and sensing integration node is treated as an agent. The collaboration method of multiple agents is designed based on a "centralized training, decentralized decision-making" architecture. Generative adversarial networks are integrated into the underlying multi-agent deep reinforcement learning algorithm to improve the estimation accuracy of Q-values and the performance of resource allocation. Background Technology
[0002] Integrated sensing and communication (ISAC) promises to facilitate the development of various emerging applications, such as smart cities, by integrating communication and sensing functions into a single platform sharing the spectrum. Compared to separate radar and communication systems, ISAC offers advantages such as high spectrum efficiency, low cost, and reduced interference. Therefore, ISAC has broad application potential. However, single static ISAC systems face challenges in terms of limited service coverage and sensing capabilities due to their limited observation angle and congestion.
[0003] Fortunately, multi-user cooperative communication and sensing integrated systems based on distributed base stations can effectively expand the service range and improve sensing accuracy, and therefore have attracted widespread attention. For example, the paper [Yu Cao and Qi-Yue Yu. Joint Resource Allocation for User-Centric Cell-Free Integrated Sensing and Communication Systems. IEEE Commun. Lett., 27(9):2338–2342, Sept. 2023.] studies the resource allocation problem of communication-centric decellularized ISAC, and maximizes the system's reachability and rate by jointly designing user scheduling and power allocation. The paper [Weihao Mao, Yang Lu, Chong-Yung Chi, Bo Ai, ZhangduiZhong, and Zhiguo Ding. Communication-Sensing Region for Cell-Free MassiveMIMO ISAC Systems. IEEE Trans. Wireless Commun., 23(9):12396–12411, Sept.2024.] establishes a new performance index and studies the beamforming optimization problem of decellularized ISAC. However, the above schemes are based on traditional optimization algorithms and have challenges such as high computational complexity and poor robustness, especially in large-scale decellularization of ISAC in dynamic environments.
[0004] On the other hand, deep reinforcement learning (DRL) has recently emerged as a promising technique for handling multi-step decision-making in dynamic environments. Compared to traditional optimization algorithms, DRL can obtain the optimal policy through trial and error, offering advantages such as low computational complexity and robustness. Furthermore, the training labels for DRL are obtained through the interaction between the agent and the environment, making it widely applicable in wireless communication where samples are scarce. Compared to single-agent DRL, multi-agent DRL allows each agent to make decisions based on its own observations without requiring frequent communication between agents. Moreover, by designing rewards individually for each agent, multi-agent deep reinforcement learning can facilitate more flexible settings and multi-objective optimization. Therefore, multi-agent deep reinforcement learning is becoming a promising technique for handling distributed optimization problems. For example, the paper [Haocheng Zhang, Rang Liu, Ming Li, Wei Wang, and Qian Liu. Joint Sensing and Communication Optimization in Target-Mounted STARS-Assisted Vehicular Networks: A MADRL Approach. IEEE Trans. Veh. Tech., 73(7):10011–10025, Jul. 2024.] proposes a multi-agent DRL framework for simultaneously reflecting and transmitting surfaces to assist in ISAC, in order to achieve distributed control and multi-objective optimization.
[0005] Despite the aforementioned advantages of multi-agent deep reinforcement learning (DRL), it still faces challenges such as low sampling efficiency and training instability. Fortunately, generative artificial intelligence (GAI) has demonstrated significant capabilities in digital content generation and data analysis, and holds promise for enhancing DRL from multiple perspectives. On the one hand, GAI can generate a large number of experience tuples, bypassing the continuous interaction between the agent and the environment; on the other hand, GAI has great potential in improving the training stability and performance of DRL. Therefore, deep reinforcement learning based on GAI has attracted extensive research from numerous researchers.
[0006] This invention leverages the numerous advantages of GAI, combining GAI and multi-agent DRL to propose a generative multi-agent reinforcement learning algorithm to solve the dynamic resource allocation problem in decellularized ISAC systems. This invention maximizes the sum of signal-to-interference-plus-noise ratios (SINRs) within the F time slot, while satisfying constraints on communication service quality, transmit power, base station scheduling, and receive power. Summary of the Invention
[0007] To address the problems existing in current technologies, this invention proposes a design method for a multi-agent integrated sensing system, capable of solving the problem of maximizing the sensing signal-to-interference-plus-noise ratio in a decellularized multi-site integrated sensing system. In the network model of this invention, Several ISAC base stations cooperate to detect a moving point target and provide services. Multiple antenna users, among which and Assume that each ISAC base station has the following number of antennas. The number of antennas per user is The specific plan is shown in the diagram. Figure 1 As shown in the figure. Based on this model, the present invention provides a design method for joint transmit beamforming, receive filtering, and base station scheduling to maximize the sum of the signal-to-interference-plus-noise ratio (SINR) perceived by the system in multiple time slots.
[0008] To achieve the above-mentioned technical objectives, the technical solution adopted by the present invention is as follows: A design method for a multi-agent integrated sensing system is disclosed. This method expands the range of communication sensing services by coordinating the transmission and reception of multiple integrated base stations. Based on a "centralized training, decentralized decision-making" architecture, the method designs the state space, action space, reward function, and cooperation mode of multiple agents. To improve Q-value estimation accuracy and decision quality, a generative adversarial network is integrated into the multi-agent deep reinforcement learning method. Specifically, the method includes the following steps: The first step is to build a system model, specifically: Step 1.1, Multiple antenna ISAC base stations work together to detect a moving point target and simultaneously provide services. Multiple antenna users. Assume the number of antennas per ISAC base station and per user are respectively... , ,in and In the first The first time slot, the first From the first ISAC base station to the mobile target, the first The number of ISAC base stations to the first The user, the The number of ISAC base stations to the first The channels of each base station are represented as follows: , and ; Step 1.2, in order to strike a balance between performance and complexity, the system only activates [specific functions] at a time. There are 1 ISAC base stations, of which The echoes from all activated ISAC base stations are uploaded to the central server via the feedforward channel for data fusion and target detection. Step 1.3, the The ISAC base station in the first Transmission signal of time slot Represented as: (1) in, This represents the base station selection variable. , Indicates the first j The first ISAC-BS is selected; otherwise, the second... j One ISAC base station was not selected; Represents the communication beamforming matrix; Represents a communication symbol vector; Represents the sensing beamforming vector; Representing perceptual symbols; Represents the communication sensing beamforming matrix; Represents a communication-aware symbol vector; Indicates the ISAC base station index; This represents the time slot index. Specifically: and Representing communication symbol vectors and sensing signals, satisfying , and . and Corresponding to and The transmitted beamforming matrix. For convenience, define... and ; No. In the first time slot k The received signal for each user is represented as follows: (2) in, Indicates user k The received signal; j Indicates the index of the ISAC base station. JIndicates the number of ISAC base stations; Indicates the user index; This represents additive white Gaussian noise. Indicates noise power. Represents the identity matrix; Indicates the first In the first time slot The number of ISAC base stations to the first k Channels for individual users; Without loss of generality, channel Specifically, the Rice channel has the following expression: (3) in, This represents the path loss at a reference distance of 1 meter; Indicates the path loss index; Indicates the first The ISAC base station and the first k The distance between individual users; Represents Rice factor; This refers to the channel in the non-line-of-sight portion; This represents the line-of-sight portion of the channel, where, Indicates the first k Number of antennas per user This indicates the number of transmit antennas in an ISAC base station. Indicates the first k The user relative to the first Azimuth angle of each ISAC base station Indicates the first The ISAC base station relative to the first k The azimuth angle of each user. Indicates the array steering vector. Represents the conjugate transpose of the array guide vector; specifically: Represent a N A uniform linear array of elements at an angle Directional guide vector, Indicates azimuth. Indicates the number of antennas.
[0009] also, Represented as: (4) in, , , The first j The distance, path loss index, and azimuth of each ISAC base station to the target; This represents the path loss at a reference distance of 1 meter. Indicates the imaginary part.
[0010] Assuming each user performs receive filtering on the received signal using a linear filter to improve its signal-to-noise ratio, then the... k The user in the first Output signal-to-noise ratio of each time slot Represented as: (5) in, Indicates the first j Selection variables for each ISAC base station; Indicates noise power; Represents the square of the L2 norm; j Indicates the index of the ISAC base station; k Indicates the user index; J Indicates the number of ISAC base stations; K Indicates the number of users; Indicates the first k The coefficients of a communication receiver filter; express The k List; No. j The ISAC base station in the first The received signal for each time slot is represented as follows: (6) in, Indicates the first j The received signals of each base station; Indicates the first Channels from each ISAC base station to the target; express The conjugate transpose of; The radar cross-section of the target is given by its power. ; This represents additive white Gaussian noise with a noise power of . ; It refers to the first In the first time slot The number of ISAC base stations to the first The direct-fire channel of each ISAC base station is in the following form: (7) in, This is the path loss index. ; For the first i The ISAC base station and the first j The distance between ISAC base stations; For the first j The ISAC base station and the first i Azimuth angle of each ISAC base station; By uploading the echo data from all activated ISAC base stations to the central server, the fused signal can be obtained. (8) in: Indicates a fused signal; This indicates the received signal from the first ISAC base station; Indicates the first J Received signals from one ISAC base station; Represents the equivalent target response matrix; Represents the equivalent transmitted signal vector; This represents the response matrix between equivalent base stations; Represents an equivalent noise vector; specifically: (9) in: This represents the equivalent selection variable from ISAC base station 1 to ISAC base station 2; Indicates ISAC base station 1 to ISAC base station J Equivalent selection variables; Indicates ISAC base station J The equivalent selection variable to ISAC base station 1, Indicates ISAC base station J To ISAC base station J Equivalent selection variables; This represents the equivalent target response matrix from ISAC base station 1 to the target and back to ISAC base station 1; This indicates the route from ISAC base station 1 to the target and then back to ISAC base station. J The equivalent objective response matrix; Indicates ISAC base station J The equivalent target response matrix from the target to ISAC base station 1; Indicates ISAC base station J To the target and then to the ISAC base station J The equivalent objective response matrix; Indicates ISAC base station 1 to ISAC base station J The response matrix; Indicates ISAC base station J To ISAC base station J The equivalent response matrix; specifically: , , and ; After receiving and filtering, the central server in the [number]th [period] Output sensing signal-to-noise ratio per time slot Represented as: (10) in, Represents the mathematical expectation; These represent the coefficients of the sensing receiver filter. express The conjugate transpose of; Represents the set of all beamforming matrices; The second step is to determine the objective function and optimization variables, and to list the optimization problem: Building upon the first step, a collaborative communication sensing optimization problem involving multiple base stations is designed. Specifically, this problem aims to maximize the efficiency of communication while satisfying constraints on quality of service, transmit power, base station scheduling, and receive power. L The sum of the perceived signal-to-interference-plus-noise ratios (SIRs) for each time slot leads to the following optimization problem: (11) in, Represents the communication sensing beamforming matrix; This represents the coefficient of the k-th communication receiving filter; These represent the coefficients of the sensing receiver filter; This represents the base station selection variable. ; Indicates the time slot index; Indicates the total number of time slots; Indicates the first Output sensing signal-to-noise ratio for each time slot; Indicates the first k The user in the first Output signal-to-noise ratio per time slot; Indicates the first The signal-to-interference-plus-noise ratio threshold for individual users' communication. ; Indicates the user index; Indicates the index of the ISAC base station; J Indicates the number of ISAC base stations; Denotes the F-norm; Indicates the maximum transmission power. ; This indicates the number of active ISAC base stations; C1 represents the L2 norm; C2 represents the communication service quality constraint; C3 and C4 represent the base station scheduling constraint; and C5 and C6 represent the receive power constraint.
[0011] It can be seen that formula (11) is a non-convex mixed integer optimization problem consisting of a non-concave objective function, non-concave constraints, and coupling variables. Furthermore, dynamic environments make traditional convex / non-convex optimization algorithms difficult to handle. In this invention, we utilize Multi-Agent Deep Reinforcement Learning (MADRL) to handle dynamic resource allocation. Specifically, the cooperating ISAC base station and each user are treated as agents. To reduce overhead, each agent makes decisions based on its own observations; The third step is to design a multi-agent reinforcement learning algorithm to solve the optimization problem: First, the optimization problem shown in formula (11) is transformed into a Markov decision process, which consists of agents, environment, state, reward, state transition, etc. Each agent generates a large number of experience arrays through continuous interaction with the environment and stores them in the experience replay buffer for training agents. Then, an algorithm based on multi-agent DRL is developed to solve the problem. In particular, the decellularized ISAC system can be regarded as the environment, and each user and all ISAC base stations are regarded as agents. The actions, states and rewards of the agents are as follows: Step 3.1, Action; Each agent's action is its optimization variable. Therefore, the ISAC base station agent (the agent) K +1) in the The actions within each time slot are a set of transmitted beamforming matrices. Sensing and receiving filter coefficients and base station scheduling The action is then represented as ,in Representation matrix Vectorization, Representation matrix Vectorization, This represents the selection variable for the first ISAC base station. Indicates the first J The selection variables for each ISAC base station. Since the input to the neural network is a real number, its input is... The cascade of the real and imaginary parts, with dimension . Furthermore, the first k The user agent in the first... The action within each time slot is to receive the beamforming vector. Its dimensions are ;in Indicates the number of transmit antennas of the ISAC base station. Indicates the number of ISAC base stations. Indicates the number of users.
[0012] Step 3.2, Status; The state of an agent represents the environmental information it observes; it should include the current channel and the actions taken in the previous time slot. (ISAC base station agent (the...)) K +1 agent) in the 1st The state in each time slot should include all communication and sensing channels related to ISAC-BS, as well as the actions in the previous time slot, represented as: (12) in, This represents the state of the ISAC base station agent (the (K+1)th agent); This indicates the channel from the first ISAC base station to the mobile target; Indicates the first J Channels from an ISAC base station to a mobile target; This represents the vectorization of the channel from the first ISAC base station to the first user; Indicates the first J The number of ISAC base stations to the first K Vectorization of channels for individual users; This represents the vectorization of the channel from the first ISAC base station to the second ISAC base station; Indicates the first J Vectorization of the channel from the -1th ISAC base station to the Jth ISAC base station; This represents the vectorization of the actual beamforming matrix of the first ISAC base station; This represents the vectorization of the actual beamforming matrix of the J-th ISAC base station; This represents the vectorization of the received beamforming matrix; This represents the selection variable for the first ISAC base station; Indicates the first J Selection variables for each ISAC base station; The dimension of action is Since users cannot communicate with each other, each user agent only observes its current wireless channel and the operations in the previous time slot as its state. Therefore, the... k The state of each user agent is represented as follows: Its dimensions are , ; Step 3.3, Reward; Since different agents have different goals, the reward function is tailored to each agent. Each user agent's goal is to maximize the signal-to-interference-plus-noise ratio (SIR) of the communication; therefore, the reward function is... k The user agent in the first... Rewards within a certain time period It can be represented as: (13) in, This represents the signal-to-interference-plus-noise ratio (SIR) of the k-th user. Indicates the first k Rate threshold per user Indicates the user index; The goal of an ISAC base station is to maximize the perceived signal-to-noise ratio while meeting communication requirements. Therefore, the ISAC base station agent (the...) K +1 agent) in the 1st Rewards per time slot for: (14) in, To balance the weighting coefficients of sensing and communication performance, ; Indicates the ISAC base station agent (the first K The signal-to-interference-plus-noise ratio (SIR) of +1 agent; This represents the signal-to-interference-plus-noise ratio (SIR) of the first agent. Indicates the first K Signal-to-interference-to-noise ratio of each agent; Step 3.4, implement a multi-agent reinforcement learning algorithm; Step 3.4.1: Leveraging its superior data analysis capabilities, this invention replaces the commentator network in the TD3 algorithm with two Generative Adversarial Networks (GANs) to improve training stability and estimation accuracy, naming them GAN-TD3. (The text then repeats the first two steps.) k The framework of the proposed GAN-TD3 algorithm is as follows: (The algorithm is described using a few intelligent agents.) No. k An agent uses an action network to determine its policy. This includes an active action network. A target action network They each have parameters and , Let represent the learnable parameters of the k-th active action network. This represents the learnable parameters of the k-th target action network. Both the main action network and the target action network have... A multilayer perceptron with layers, each hidden layer has There are 10 neurons. Furthermore, ReLU and Tanh are used as activation functions for the hidden and output layers, respectively. To facilitate distributed execution and reduce overhead, each agent inputs its own state, rather than the joint state of all agents, into its action network to obtain the corresponding action. The output of the normalized action network, as shown in equations (15) to (17), satisfies constraints C2, C5, and C6 in the optimization problem shown in equation (11); specifically, the normalization corresponding to C2 can be expressed as: (15) in, Indicates the maximum transmission power; Denotes the F-norm; The normalization corresponding to C5 can be expressed as: (16) The normalization corresponding to C6 can be expressed as: (17) Furthermore, this invention proposes a method for continuous values The mechanism for mapping to binary values satisfies constraint C3. Specifically, this invention will... Sort in descending order, then sort by the first few. Set one element to 1 and the rest to 0; Step 3.4.2, the purpose of the generator is to estimate the Q-value. Furthermore, the generator consists of two main generators. and Composition, each of which has parameters and There are also two target generators. and Each has parameters and ,in, Indicates the first primary generator. Indicates the second primary generator. This represents the parameters of the first primary generator. This represents the parameters of the second primary generator. Indicates the first target generator. This represents the second target generator. This represents the parameters of the first target generator. This represents the parameters of the second target generator. Each generator network is a network with... A multi-sensor with layers, each hidden layer has The generator network has 10 neurons and uses ReLU as the activation function. In addition to states and actions, it also takes a one-dimensional Gaussian random variable as input. Step 3.4.3, the discriminator focuses on distinguishing between the true and estimated Q values, and includes two discriminators. and They have parameters and ,in, This represents the first discriminator. This represents the second discriminator. Indicates the parameters of the first discriminator. This represents the parameters of the second discriminator. The discriminator architecture is a... one neuron A multilayer perceptron. Furthermore, the discriminator outputs the probability that the input Q value is the true value. Note that... and , and These are two generative adversarial networks designed to improve the accuracy of Q-value estimation; Step 3.4.4: Once enough experience tuples have been collected, the training process begins by randomly selecting a batch of experience tuples from the replay buffer to train the actor and critic networks. k On the first intelligent agent n The loss function for a generator can be written as: (18) in, Indicates the first k On the first intelligent agent n Loss function of each generator; Indicates the first k On the first intelligent agent n The output of each generator; Indicates state; Indicates an action; Represents the training sample set; Indicates the first k On the first intelligent agent n The output of the discriminator; This represents the target Q value, specifically: (19) in, Indicates the first k One reward; Indicates the discount factor; Indicates the first k On the first intelligent agent n The output of the target generator; express The state at any given moment; Representation strategy; No. The loss of the discriminator It can be written as: (20) in, Indicates the first k On the first intelligent agent n The output of the discriminator; Indicates the first k On the first intelligent agent n The output of each generator; Update the parameters of the master generator and discriminator using the gradient descent algorithm: (twenty one) in, Indicates coefficient; Indicates coefficient; Relative to The gradient; Relative to The gradient; Step 3.4.5: Additionally, the parameter update method for the target generator is the same as that for the target commentator network.
[0013] Action networks update their parameters using the policy gradient method: (twenty two) in, Indicates the learning rate. ; Represents the parameters of the action network; Indicates about The gradient; Indicates about The gradient; Indicates about The gradient; The parameter is The strategy.
[0014] In addition, the parameters of the target action network are updated using a soft update method: (twenty three) In the formula, Represents the coefficient. ; Indicates the previous parameter; This indicates a new parameter.
[0015] Step 3.5: Solve the optimization problem shown in formula (11) above through continuous interaction between the agent and the environment. In each training round, each agent interacts with the environment to obtain an experience replay array and puts it into the experience replay buffer. Sample the experience replay array from the experience replay buffer, update the network parameters, and perform the next training round until the maximum number of training rounds is reached. The specific process is as follows: Step 3.5.1, Set the initial state vector The parameter values of the action network, generator, and target generator for each agent. , and ; Step 3.5.2, let , , ; Step 3.5.3: Each agent interacts with the environment to obtain an experience replay array and stores it in the experience replay buffer pool; Step 3.5.4: Randomly sample a batch of samples from the experience replay buffer pool; Step 3.5.5, update according to formula (21) , , and ; Step 3.5.6, update according to formula (22) Update using soft update method , and ; Step 3.5.7: If the maximum number of training iterations has not been reached, return to step 3.5.3; otherwise, the algorithm training ends.
[0016] This invention considers a design method for a multi-agent integrated sensing system. In order to maximize the sum of the signal-to-interference-plus-noise ratio of multiple time slot sensing echoes, the transmit beamforming, receive filtering and base station scheduling are jointly optimized. The multi-agent reinforcement learning algorithm ensures the cooperative performance of the multi-station integrated sensing system and reduces the computational complexity.
[0017] The beneficial effects of this invention are: (1) This invention provides a scheme to maximize the sum of signal-to-interference-plus-noise ratio (SNR) of multi-slot sensing echoes under the conditions of satisfying transmit power and base station scheduling by jointly optimizing transmit beamforming, receive filtering and base station scheduling of multiple base stations. This invention is based on a reinforcement learning algorithm enhanced by generative artificial intelligence, which transforms the original problem into a Markov decision process, and learns the optimal resource allocation and multi-station cooperation strategy through the interaction between the agent and the environment.
[0018] (2) To improve the performance of multi-agent reinforcement learning algorithms, this invention integrates generative adversarial networks into the multi-agent reinforcement learning algorithm, thereby improving the accuracy of Q-value estimation and the quality of decision-making. This invention provides a reference value method for designing a high-performance, robust, multi-station communication and sensing integrated system. The designed multi-agent reinforcement learning algorithm has advantages such as strong generalization and low complexity. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of a multi-station communication and sensing integrated network.
[0020] Figure 2 It is the GANTD3 algorithm framework.
[0021] Figure 3 These are the convergence curves of different algorithms; Figure 3 (a) in the figure represents the convergence curve of user agent 1 under different algorithms; Figure 3 (b) in the figure represents the convergence curve of user agent 2 under different algorithms; Figure 3 (c) in the figure represents the convergence curve of user agent 3 under different algorithms; Figure 3 In the figure, (d) represents the convergence curve of the integrated sensing base station agent (agent 4) under different algorithms.
[0022] Figure 4 It is the curve of agent reward as a function of transmission power; Figure 4 In the figure, (a) shows the curve of reward for user agent 1 as a function of transmission power under different algorithms; Figure 4 (b) in the figure shows the curves of reward for user agent 2 as a function of transmission power under different algorithms; Figure 4 (c) in the figure represents the curve of reward for user agent 3 as a function of transmission power under different algorithms; Figure 4 In the figure, (d) represents the curve of reward change with transmission power for the synesthetic agent (agent 4) under different algorithms.
[0023] Figure 5 It is a curve showing how the agent's reward changes with the number of transmitting antennas; Figure 5 In the figure, (a) shows the curve of the reward of user agent 1 as a function of the number of transmitting antennas under different algorithms; Figure 5 (b) in the figure shows the curves of the reward of user agent 2 as a function of the number of transmitting antennas under different algorithms; Figure 5 (c) in the figure represents the curve of the reward of user agent 3 as a function of the number of transmitting antennas under different algorithms; Figure 5 In the figure, (d) is the curve of the reward of the integrated sensing base station agent (agent 4) as a function of the number of transmitting antennas under different algorithms. Detailed Implementation
[0024] To better understand the above technical solution, a detailed analysis is provided below in conjunction with the accompanying drawings and specific implementation methods.
[0025] A method for designing a multi-agent synesthetic system includes the following steps: The first step is to build a system model: Multiple antenna ISAC base stations work together to detect a moving point target and provide services. Multiple antenna users. Assume the number of antennas per ISAC base station and per user are respectively... and To strike a balance between performance and complexity, only activate... There are 15 ISAC base stations. The echoes from all active ISAC base stations are uploaded to the central server via the feedforward channel for data fusion and target detection.
[0026] The second step is to determine the objective function and optimization variables, and to list the optimization problem: By jointly optimizing the transmit beamforming, receive filtering, and base station scheduling of multiple base stations, a scheme is proposed to maximize the sum of the signal-to-interference-plus-noise ratio (SNR) of multi-timeslot sensing echoes under the conditions of satisfying transmit power and base station scheduling. Based on a reinforcement learning algorithm enhanced by generative artificial intelligence, the original problem is transformed into a Markov decision process, in which the agent learns the optimal resource allocation and multi-station cooperation strategy through interaction with the environment. The third step is to design a multi-agent reinforcement learning algorithm to solve the optimization problem: (1) Definition of state space; Define the state space for each agent as shown in Equation (12); (2) Definition of action space; Define the action space for each agent; (3) Setting the reward function; Define the reward function for each agent as shown in formulas (13) and (14); The fourth step is to design a reinforcement learning algorithm: Based on the optimization problem studied, a generative adversarial network-enhanced TD3 algorithm is proposed, and a normalized operator for the action network output is designed to satisfy some constraints. Each agent is trained based on a "centralized training, decentralized decision-making" architecture, as detailed below: 1) Set the initial state vector The parameter values of the action network, generator, and target generator for each agent. , and ; 2) Order , , ; 3) Each agent interacts with the environment to obtain an experience replay array and stores it in the experience replay buffer pool; 3) Randomly sample a batch of samples from the experience playback buffer; 4) Update according to formula (21) , , and ; 5) Update according to formula (22) Update using soft update method , and ; 6) If the maximum number of training iterations has not been reached, return to step 3; otherwise, the algorithm training ends.
[0027] This embodiment verifies: and The positions of ISAC-BS are set in metric units. , , and The user's coordinates are set in metric units. , and The target starts from its initial position. Begin moving, with a movement speed of [number] times per time slot. rice.
[0028] (1) The method proposed in this embodiment adopts the generative multi-agent reinforcement learning algorithm (MAGANTD3) and considers two benchmarks for comparison: the first is the multi-agent deep deterministic policy gradient algorithm (MADDPG); the second is the multi-agent twin-delayed deep deterministic policy gradient algorithm (MATD3).
[0029] (2) Analyze the convergence curves of different schemes: Figure 3 The reward curves of four agents under different multi-agent reinforcement learning algorithms are shown. Figure 3 In (a), user agent 1 converges within 3000 rounds under all three multi-agent reinforcement learning methods, and the MAGANTD3 algorithm outperforms the other two. Figure 3 In (b), user agent 2 converges within 3000 rounds under all three multi-agent reinforcement learning methods, and the MAGANTD3 algorithm outperforms the other two. Figure 3 In (c), user agent 3 converges within 3000 rounds under all three multi-agent reinforcement learning methods, and the MAGANTD3 algorithm outperforms the other two. Figure 3 In (d), the integrated sensory base station agent (agent 4) can converge within 3000 rounds under the three multi-agent reinforcement learning methods, and the MAGANTD3 algorithm is superior to the other two algorithms.
[0030] (3) Analyze the curve of agent reward as a function of transmission power: Figure 4 It demonstrates the relationship between the agent's reward and transmission power.
[0031] exist Figure 4 In (a), user agent 1 achieves similar performance under MAGANTD3 and MATD3 and outperforms the MADDPG algorithm. Figure 4 In (b), user agent 2 achieves similar performance under MAGANTD3 and MATD3 and outperforms the MADDPG algorithm. Figure 4 In (c), user agent 3 achieves similar performance under MAGANTD3 and MATD3 and outperforms the MADDPG algorithm. Figure 4 In (d), the integrated sensing base station agent (agent 4) achieves similar performance under MAGANTD3 and MATD3 and outperforms the MADDPG algorithm.
[0032] (4) Analyze the curve of agent reward as a function of the number of transmitting antennas: Figure 5 It demonstrates the relationship between the agent's reward and the number of transmitting antennas.
[0033] exist Figure 5 In (a), user agent 1's reward under MAGANTD3 and MATD3 almost always increases with the increase of the number of transmit antennas and outperforms the MADDPG algorithm due to the increase in degrees of freedom. However, MADDPG struggles with learning high-quality policies, especially in cases with a large action space. Figure 5 In (b), user agent 2's reward under MAGANTD3 and MATD3 almost always increases with the increase of the number of transmit antennas and outperforms the MADDPG algorithm. Figure 5In (c), the reward for user agent 3 under MAGANTD3 and MATD3 almost always increases with the increase of the number of transmit antennas and outperforms the MADDPG algorithm. Figure 5 In (d), the reward of the integrated sensing base station agent (agent 4) under MAGANTD3 and MATD3 increases with the number of transmitting antennas and is better than the MADDPG algorithm.
[0034] The above-described embodiments are merely illustrative of the implementation methods of the present invention, but should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the protection scope of the present invention.
Claims
1. A method for designing a multi-agent synesthetic integrated system, characterized in that, The method for designing a multi-agent synesthetic integrated system includes the following steps: The first step is to build a system model; The second step is to determine the objective function and optimization variables, and to list the optimization problem: Design a perception optimization problem for collaborative communication among multiple base stations; maximize the following while satisfying constraints on communication service quality, transmit power, base station scheduling, and receive power: L The sum of the perceived signal-to-interference-plus-noise ratios (SIRs) for each time slot leads to the following optimization problem: (11) in, Represents the communication sensing beamforming matrix; This represents the coefficient of the k-th communication receiving filter; These represent the coefficients of the sensing receiver filter; Indicates the base station selection variable. ; Indicates the time slot index; Indicates the total number of time slots; Indicates the first Output sensing signal-to-noise ratio for each time slot; Indicates the first k The user in the first Output signal-to-noise ratio per time slot; Indicates the first The signal-to-interference-plus-noise ratio threshold for individual users' communication. ; Indicates the user index; Indicates the index of the ISAC base station; J Indicates the number of ISAC base stations; Denotes the F-norm; Indicates the maximum transmission power. ; This indicates the number of active ISAC base stations; C1 represents the L2 norm; C2 represents the communication service quality constraint; C3 and C4 represent the base station scheduling constraint; and C5 and C6 represent the receive power constraint. The third step is to design a multi-agent reinforcement learning algorithm to solve the optimization problem: The optimization problem shown in formula (11) is transformed into a Markov decision process, which consists of an agent, environment, state, reward, and state transition. Each agent generates a large number of experience arrays through continuous interaction with the environment and stores them in the experience replay buffer for training agents. An algorithm based on multi-agent DRL is designed to solve the problem.
2. The method for designing a multi-agent synesthetic integrated system according to claim 1, characterized in that, The first step is specifically as follows: Step 1.1, Multiple antenna ISAC base stations work together to detect a moving point target and simultaneously provide services. Multiple antenna users; assuming the number of antennas for each ISAC base station and each user are respectively... , ,in and ; in the The first time slot, the first From the first ISAC base station to the mobile target, the first The number of ISAC base stations to the first The user, the first The number of ISAC base stations to the first The channels of each base station are represented as follows: , and ; Step 1.2, to activate only the system at a time There are 1 ISAC base stations, of which The echoes from all activated ISAC base stations are uploaded to the central server via the feedforward channel for data fusion and target detection. Step 1.3, the The ISAC base station in the first Transmission signal of time slot Represented as: (1) in, Indicates the base station selection variable. , Indicates the first j The first ISAC-BS is selected; otherwise, the second... j One ISAC base station was not selected; Represents the communication beamforming matrix; Represents a communication symbol vector; Represents the sensing beamforming vector; Representing perceptual symbols; Represents the communication sensing beamforming matrix; Represents a communication-aware symbol vector; Indicates the ISAC base station index; Represents the time slot index; specifically: and Representing communication symbol vectors and sensing signals, satisfying , and ; and Corresponding to and The transmitted beamforming matrix; definition and ; No. In the first time slot k The received signal for each user is represented as follows: (2) in, Indicates user k The received signal; j Indicates the index of the ISAC base station. J Indicates the number of ISAC base stations; Indicates the user index; This represents additive white Gaussian noise. Indicates noise power. Represents the identity matrix; Indicates the first In the first time slot The number of ISAC base stations to the first k Channels for individual users; Assuming each user performs receive filtering on the received signal using a linear filter to improve its signal-to-noise ratio, then the... k The user in the first Output signal-to-noise ratio of each time slot Represented as: (5) in, Indicates the first j Selection variables for each ISAC base station; Indicates noise power; Represents the square of the L2 norm; j Indicates the index of the ISAC base station; k Indicates the user index; J Indicates the number of ISAC base stations; K Indicates the number of users; Indicates the first k The coefficients of a communication receiver filter; express The k List; No. j The ISAC base station in the first The received signal for each time slot is represented as follows: (6) in, Indicates the first j The received signals of each base station; Indicates the first Channels from each ISAC base station to the target; express The conjugate transpose of; The radar cross-section of the target is given by its power. ; This represents additive white Gaussian noise with a noise power of . ; It refers to the first In the first time slot The number of ISAC base stations to the first The direct-fire channel of each ISAC base station is in the following form: (7) in, This is the path loss index. ; For the first i The ISAC base station and the first j The distance between ISAC base stations; For the first j The ISAC base station and the first i Azimuth angle of each ISAC base station; The echo data from all activated ISAC base stations are uploaded to the central server to obtain the fused signal. (8) in: Indicates a fused signal; This indicates the received signal from the first ISAC base station; Indicates the first J Received signals from one ISAC base station; Represents the equivalent target response matrix; Represents the equivalent transmitted signal vector; This represents the response matrix between equivalent base stations; Represents an equivalent noise vector; After receiving and filtering, the central server in the [number]th [period] Output sensing signal-to-noise ratio per time slot Represented as: (10) in, Represents the mathematical expectation; These represent the coefficients of the sensing receiver filter. express The conjugate transpose of; This represents the set of all beamforming matrices.
3. The method for designing a multi-agent synesthetic integrated system according to claim 2, characterized in that, In step 1.3: The channel Specifically, it is a Rice channel, and its expression is: (3) in, This represents the path loss at a reference distance of 1 meter; Indicates the path loss index; Indicates the first The ISAC base station and the first k The distance between individual users; Represents Rice factor; This refers to the channel in the non-line-of-sight portion; This represents the line-of-sight portion of the channel, where, Indicates the first k Number of antennas per user This indicates the number of transmit antennas in an ISAC base station. Indicates the first k The user relative to the first Azimuth angle of each ISAC base station Indicates the first The ISAC base station relative to the first k The azimuth angle of each user. Indicates the array steering vector. This represents the conjugate transpose of the array guide vector; specifically: Represent a N A uniform linear array of elements at an angle Directional guide vector, Indicates azimuth. Indicates the number of antennas; also, Represented as: (4) in, , , The first j The distance, path loss index, and azimuth of each ISAC base station to the target; This represents the path loss at a reference distance of 1 meter. Indicates the imaginary part unit; The equivalent target response matrix Response matrix between equivalent base stations Specifically: (9) in: This represents the equivalent selection variable from ISAC base station 1 to ISAC base station 2; Indicates ISAC base station 1 to ISAC base station J Equivalent selection variables; Indicates ISAC base station J The equivalent selection variable to ISAC base station 1, Indicates ISAC base station J To ISAC base station J Equivalent selection variables; This represents the equivalent target response matrix from ISAC base station 1 to the target and back to ISAC base station 1; This indicates the route from ISAC base station 1 to the target and then back to ISAC base station. J The equivalent objective response matrix; Indicates ISAC base station J The equivalent target response matrix from the target to ISAC base station 1; Indicates ISAC base station J To the target and then to the ISAC base station J The equivalent objective response matrix; Indicates ISAC base station 1 to ISAC base station J The response matrix; Indicates ISAC base station J To ISAC base station J The equivalent response matrix; specifically: , , and .
4. The method for designing a multi-agent synesthetic integrated system according to claim 3, characterized in that, In the second step, for the non-convex mixed integer optimization problem shown in formula (11), multi-agent deep reinforcement learning (MADRL) is used to handle dynamic resource allocation; The cooperating ISAC base stations and each user are treated as intelligent agents.
5. The method for designing a multi-agent synesthetic integrated system according to claim 4, characterized in that, In the third step, the decellularized ISAC system is considered as the environment, and each user and all ISAC base stations are treated as an agent. The actions, states, and rewards of the agents are as follows: Step 3.1, Action; The action of each agent is its optimization variable; ISAC base station intelligent agents have a total of K +1 agent, then the ISAC base station agent in the th... The actions within each time slot are a set of transmitted beamforming matrices. Sensing and receiving filter coefficients and base station scheduling The action is then represented as ,in Representation matrix Vectorization, Representation matrix Vectorization, This represents the selection variable for the first ISAC base station. Indicates the first J Selection variables for each ISAC base station; The input to a neural network is a real number; its input is... The cascade of the real and imaginary parts, with dimension . ;No. k The user agent in the first... The action within each time slot is to receive the beamforming vector. Its dimensions are ;in Indicates the number of transmit antennas of the ISAC base station. Indicates the number of ISAC base stations. Indicates the number of users; Step 3.2, Status; The agent's state represents observed environmental information, including the current channel and actions taken in the previous time slot; the ISAC base station agent in the... The state in each time slot should include all communication and sensing channels related to ISAC-BS, as well as the actions in the previous time slot; The dimension of action is ;No. k The state of each user agent is represented as follows: Its dimensions are , ; Step 3.3, Reward; The goal of each user agent is to maximize the signal-to-interference-plus-noise ratio (SIR) in communication. k The user agent in the first... Rewards within a certain time period Represented as: (13) in, This represents the signal-to-interference-plus-noise ratio (SIR) of the k-th user. Indicates the first k Rate threshold per user Indicates the user index; The ISAC base station agent in the... Rewards per time slot for: (14) in, To balance the weighting coefficients of sensing and communication performance, ; Indicates the signal-to-interference-plus-noise ratio (SIRR) of the ISAC base station agent; This represents the signal-to-interference-plus-noise ratio (SIR) of the first agent. Indicates the first K Signal-to-interference-to-noise ratio of each agent; Step 3.4, implement a multi-agent reinforcement learning algorithm; Step 3.4.1: Replace the commentator network in the TD3 algorithm with two generative adversarial networks (GANs), and name it the GAN-TD3 algorithm; Step 3.4.2, the purpose of the generator is to estimate the Q value; the generator consists of two main generators. and Composition, each with parameters and There are also two target generators. and Each has parameters and ,in, Indicates the first primary generator. Indicates the second primary generator. This represents the parameters of the first primary generator. This represents the parameters of the second primary generator. Indicates the first target generator. This represents the second target generator. This represents the parameters of the first target generator. This represents the parameters of the second target generator; each generator network is a network with... A multi-sensor with layers, each hidden layer has The generator network has 100 neurons and uses ReLU as the activation function; in addition to state and action, it also takes a one-dimensional Gaussian random variable as input. Step 3.4.3: The discriminator is used to distinguish between the true and estimated Q values, and includes two discriminators. and , has parameters and ,in, This represents the first discriminator. This represents the second discriminator. Indicates the parameters of the first discriminator. This represents the parameters of the second discriminator; the architecture of the discriminator is a... one neuron A multilayer perceptron; in addition, the discriminator outputs the probability that the input Q value is the true value; Step 3.4.4: Once the experience tuples are collected, the training process begins. A batch of experience tuples is randomly selected from the replay buffer to train the actor and critic networks; k On the first intelligent agent n The loss function for a generator can be written as: (18) in, Indicates the first k On the first intelligent agent n Loss function of each generator; Indicates the first k On the first intelligent agent n The output of each generator; Indicates state; Indicates an action; Represents the training sample set; Indicates the first k On the first intelligent agent n The output of the discriminator; This represents the target Q value, specifically: (19) in, Indicates the first k One reward; Indicates the discount factor; Indicates the first k On the first intelligent agent n The output of the target generator; express The state at any given moment; Representation strategy; No. The loss of the discriminator Written as: (20) in, Indicates the first k On the first intelligent agent n The output of the discriminator; Indicates the first k On the first intelligent agent n The output of each generator; Update the parameters of the master generator and discriminator using the gradient descent algorithm; Step 3.4.5: The parameter update method for the target generator is the same as that for the target critic network; The action network updates its parameters using the policy gradient method; the parameters of the target action network are updated using a soft update method. Step 3.5: Solve the optimization problem shown in the above formula (11) through continuous interaction between the agent and the environment. In each round of training, each agent interacts with the environment to obtain the experience replay array and puts it into the experience replay buffer pool. Sample the experience replay array from the experience replay buffer pool, update the network parameters, and perform the next training until the maximum number of training times is reached.
6. The method for designing a multi-agent synesthetic integrated system according to claim 5, characterized in that, In step 3.2, the state is represented as follows: (12) in, Indicates the state of the ISAC base station agent; This indicates the channel from the first ISAC base station to the mobile target; Indicates the first J Channels from an ISAC base station to a mobile target; This represents the vectorization of the channel from the first ISAC base station to the first user; Indicates the first J The number of ISAC base stations to the first K Vectorization of channels for individual users; This represents the vectorization of the channel from the first ISAC base station to the second ISAC base station; Indicates the first J Vectorization of the channel from the -1th ISAC base station to the Jth ISAC base station; This represents the vectorization of the actual beamforming matrix of the first ISAC base station; This represents the vectorization of the actual beamforming matrix of the J-th ISAC base station; This represents the vectorization of the received beamforming matrix; This represents the selection variable for the first ISAC base station; Indicates the first J The selection variables for each ISAC base station.
7. The method for designing a multi-agent sensory integrated system according to claim 6, characterized in that, In step 3.4.1, the first k The framework of the proposed GAN-TD3 algorithm is as follows: (The algorithm is described using a few intelligent agents.) No. k An agent uses an action network to determine its policy. This includes an active action network. A target action network Each has parameters and , Let represent the learnable parameters of the k-th active action network. This represents the learnable parameters of the k-th target action network; both the main action network and the target action network have... A multilayer perceptron with layers, each hidden layer has There are 10 neurons; and ReLU and Tanh are used as activation functions for the hidden and output layers, respectively. Each agent inputs its own state, rather than the joint state of all agents, into its action network to obtain the corresponding action; the output of the normalized action network, as shown in formulas (15) to (17), satisfies the constraints C2, C5, and C6 in the optimization problem shown in formula (11), and C2, C5, and C6 are normalized respectively; in addition, a method for continuous values is proposed. The mechanism for mapping to binary values satisfies constraint C3.
8. The method for designing a multi-agent sensory integrated system according to claim 7, characterized in that, The normalization process for C2, C5, and C6 is as follows: The normalized representation corresponding to C2 is: (15) in, Indicates the maximum transmission power; Denotes the F-norm; The normalized representation of C5 is: (16) The normalized representation of C6 is: (17) And Sort in descending order, then sort by the first few. Set one element to 1 and the rest to 0.
9. The method for designing a multi-agent sensory integrated system according to claim 8, characterized in that, In step 3.4.4, the parameters of the main generator and discriminator are updated as follows: (21) in, Indicates coefficient; Indicates coefficient; Relative to The gradient; Relative to The gradient; In step 3.4.5, the action network updates its parameters using the policy gradient method, specifically as follows: (22) in, Indicates the learning rate. ; Represents the parameters of the action network; Indicates about The gradient; Indicates about The gradient; Indicates about The gradient; The parameter is Strategies; In step 3.4.5, updating the parameters of the target action network specifically involves: (23) In the formula, Represents the coefficient. ; Indicates the previous parameter; This indicates a new parameter.
10. The method for designing a multi-agent synesthetic integrated system according to claim 9, characterized in that, Step 3.5 is as follows: Step 3.5.1, Set the initial state vector The parameter values of the action network, generator, and target generator for each agent. , and ; Step 3.5.2, let , , ; Step 3.5.3: Each agent interacts with the environment to obtain an experience replay array and stores it in the experience replay buffer pool; Step 3.5.4: Randomly sample a batch of samples from the experience replay buffer pool; Step 3.5.5, update according to formula (21) , , and ; Step 3.5.6, update according to formula (22) Update using soft update method , and ; Step 3.5.7: If the maximum number of training iterations has not been reached, return to step 3.5.3; otherwise, the algorithm training ends.