Intelligent information and energy cooperative transmission deployment method and device for MIMO system
By employing the Double DQN algorithm and convolutional neural network in the MIMO system to optimize relay node location and power allocation, the problems of relay node deployment accuracy and spectrum efficiency were solved, achieving efficient information and energy coordinated transmission and improving system performance and spectrum utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HUAZHONG UNIV OF SCI & TECH
- Filing Date
- 2023-10-13
- Publication Date
- 2026-05-19
AI Technical Summary
The deployment of relay nodes for information and energy co-transmission in existing MIMO systems lacks global intelligent consideration, resulting in low location accuracy and convergence, low spectral efficiency, and serious wireless signal power divergence.
The Double DQN algorithm is used to determine the optimal relay node location. Combined with a convolutional neural network, an intelligent hybrid beamforming method is designed to optimize the location and power allocation of the relay nodes, so as to maximize the approximate average signal-to-noise ratio and spectral efficiency from the source node to the relay node and from the relay node to the destination node.
It improves the overall service performance of information and energy coordinated transmission in MIMO systems, reduces the difficulty of relay node location planning, compensates for signal propagation loss, approaches ideal spectrum efficiency, and improves the overall energy efficiency of the system.
Smart Images

Figure CN117335846B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of communications, and more specifically, relates to an intelligent information and energy coordinated transmission deployment method and apparatus for MIMO systems. Background Technology
[0002] With the development of wireless communication technology, wireless smart terminals have been widely used. In these applications, terminal devices are usually powered by batteries, and the limited battery capacity has become a key problem that urgently needs to be solved, restricting the communication capabilities of these devices. Utilizing the characteristic of wireless radio frequency signals that can both transmit signals and carry energy, wireless information and energy can be transmitted in a coordinated manner, collecting energy while transmitting information, effectively overcoming the limitations of traditional battery capacity. This technology has attracted widespread attention from academia and industry and can be widely applied in fields such as the Internet of Things, 5G, smart homes, and wearable smart devices.
[0003] Due to the random distribution and movement of terminal nodes, communication nodes for information and energy co-transmission are often far apart, exceeding their respective transmission ranges, making direct communication infeasible or requiring high transmission power. To address this, relay nodes can be used to forward wireless signals one or more times, thereby enabling signal and energy co-transmission between non-directly connected nodes. These relay nodes offer advantages such as expanded service range, low operating costs, and improved reception power for users at the boundary of information and energy co-transmission communication.
[0004] With the rapid development of communication technology and the continuous innovation of its applications, users' requirements for the quality of information and energy co-transmission are gradually increasing. Single-antenna information and energy co-transmission can no longer meet the growing user demand. Utilizing the spatial multiplexing gain provided by MIMO (Multiple Transmitter Multiple Receiver) antennas to improve channel capacity and information and energy co-transmission rate can effectively solve problems such as low transmission quality efficiency and slow speed, and greatly promote the development of information and energy co-transmission.
[0005] Therefore, applying artificial intelligence methods, relay technology, and MIMO technology to wireless information and energy co-transmission can effectively improve the system's coverage and spectral efficiency. Current research on relays for information and energy co-transmission in MIMO systems mainly focuses on the deployment of optimal relay nodes. Due to a lack of global intelligent considerations, existing relay selection algorithms suffer from poor performance in terms of accuracy and convergence in determining relay node locations. Simultaneously, wireless signals exhibit power divergence during propagation; excessive signal divergence in useless directions reduces the spectral efficiency at the receiver, resulting in low overall system energy efficiency. Summary of the Invention
[0006] In view of the above-mentioned defects or improvement needs of the existing technology, the present invention provides an intelligent information and energy coordinated transmission deployment method and apparatus for MIMO systems, the purpose of which is to optimize the accuracy, convergence and spectral efficiency of information and energy coordinated transmission relay deployment in MIMO systems.
[0007] To achieve the above objectives, according to a first aspect of the present invention, an intelligent information and energy coordinated transmission deployment method for MIMO systems is provided, comprising:
[0008] S1, with the goal of maximizing the approximate average signal-to-noise ratio from the source node to the relay node and from the relay node to the destination node in the MIMO system, uses the Double DQN algorithm to determine the optimal relay node location, and determines the optimal relay node based on this location.
[0009] Among them, S t ={S v S r S n}、A t ={A d A l A r A s The approximate average signal-to-noise ratios from the source node to the relay node and from the relay node to the destination node at time t are respectively used as the state set, action set, and reward value at time t; S v This is a set of states for each transmitter, including transmitter power, number of transmit antennas, and transmitter system parameter states; S p This is a set of states for each receiver, including receiver received power, number of receiving antennas, and receiver system parameter states; S n A is a set of channel states, including ambient noise, channel fading, and signal-to-noise ratio; d Let A be the set of actions for a relay node to travel a straight distance Δd. l Let A be the set of actions for a relay node to turn left and travel a distance Δd. r Let A be the set of actions for a relay node to turn right and travel a distance Δd. s This is the set of actions that stop the relay node from moving.
[0010] S2, determine the optimal transmit power of the source node and the optimal forwarding power of the optimal relay node;
[0011] S3 determines the beamforming vector with the goal of maximizing the spectral efficiency of the target node.
[0012] According to a second aspect of the present invention, an intelligent information and energy coordinated transmission deployment device for MIMO systems is provided, comprising:
[0013] S1, the first processing module, is used to determine the optimal relay node location using the Double DQN algorithm with the goal of maximizing the approximate average signal-to-noise ratio from the source node to the relay node and from the relay node to the destination node in the MIMO system, and to determine the optimal relay node based on the location.
[0014] Among them, S t ={S v S r S n}、A t ={A d A l A r A s The approximate average signal-to-noise ratios from the source node to the relay node and from the relay node to the destination node at time t are respectively used as the state set, action set, and reward value at time t; S v This is a set of states for each transmitter, including transmitter power, number of transmit antennas, and transmitter system parameter states; S p This is a set of states for each receiver, including receiver received power, number of receiving antennas, and receiver system parameter states; S n A is a set of channel states, including ambient noise, channel fading, and signal-to-noise ratio; d Let A be the set of actions for a relay node to travel a straight distance Δd. l Let A be the set of actions for a relay node to turn left and travel a distance Δd. r Let A be the set of actions for a relay node to turn right and travel a distance Δd. s This is the set of actions that stop the relay node from moving.
[0015] S2, the second processing module, is used to determine the optimal transmit power of the source node and the optimal forwarding power of the optimal relay node;
[0016] S3, the third processing module, is used to determine the beamforming vector with the goal of maximizing the spectral efficiency of the target node.
[0017] According to a third aspect of the present invention, an intelligent information and energy collaborative transmission deployment system for MIMO systems is provided, comprising: a computer-readable storage medium and a processor;
[0018] The computer-readable storage medium is used to store executable instructions;
[0019] The processor is configured to read executable instructions stored in the computer-readable storage medium and execute the method as described in the first aspect.
[0020] According to a fourth aspect of the invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to perform the method as described in the first aspect.
[0021] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects:
[0022] 1. This invention combines deep reinforcement learning algorithms with a method and apparatus for deploying information and energy coordinated transmission in MIMO systems. It solves for the optimal relay node location that maximizes the approximate average signal-to-noise ratio between the source node and the relay node, and between the relay node and the destination node, and selects the optimal or near-optimal relay node based on this location. Compared to existing information and energy coordinated transmission deployment methods, this proposed method for MIMO systems effectively improves the overall service performance of information and energy coordinated transmission, reduces the difficulty of relay node location planning, overcomes the development dilemma of wireless sensor networks caused by energy shortages, and provides an effective solution for intelligent terminals to achieve information transmission and energy replenishment, thus comprehensively improving the service performance of information and energy coordinated transmission.
[0023] 2. This invention employs a convolutional neural network to design an intelligent hybrid beamforming method, which compensates for signal propagation loss in free space, effectively reduces power divergence loss during relay forwarding in existing relay systems, and can approach the ideal spectral efficiency as closely as possible. Attached Figure Description
[0024] Figure 1 This is a downlink model diagram for intelligent information and energy coordinated transmission in MIMO systems provided in an embodiment of the present invention.
[0025] Figure 2 A flowchart illustrating the intelligent information and energy coordinated transmission deployment method for MIMO systems provided in this embodiment of the invention;
[0026] Figure 3 This is a flowchart illustrating the use of the Double DQN algorithm to determine the optimal relay node location, provided as an embodiment of the present invention. Detailed Implementation
[0027] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0028] The intelligent information and energy coordinated transmission downlink model diagram for MIMO systems provided in this embodiment of the invention is shown below. Figure 1 As shown, at the start of operation, the base station first collects the original system state parameters, including the state sets of each transmitter, each receiver, and each channel. Then, deep reinforcement learning is used to select the optimal or near-optimal relay node based on each state set. Finally, the base station transmits information and energy to the target user via relay. Furthermore, the base station, relay nodes, and users all employ a MIMO system, and convolutional neural networks are used to implement intelligent hybrid beamforming with multiple antennas, compensating for signal propagation loss in free space and maximizing the spectral efficiency of the target node. Using the intelligent information and energy coordinated transmission deployment method for MIMO systems proposed in this invention, optimal performance information and energy coordinated transmission can be achieved. Specifically, as... Figure 1 As shown, this embodiment of the invention provides an intelligent information and energy collaborative transmission deployment method for MIMO systems, including:
[0029] S1, with the goal of maximizing the approximate average signal-to-noise ratio from the source node to the relay node and from the relay node to the destination node in the MIMO system, uses the Double DQN algorithm to determine the optimal relay node location, and determines the optimal relay node based on this location.
[0030] Among them, S t ={S v S r S n}、A t ={A d A l A r A s The approximate average signal-to-noise ratios from the source node to the relay node and from the relay node to the destination node at time t are respectively used as the state set, action set, and reward value at time t; S v This is a set of states for each transmitter, including transmitter power, number of transmit antennas, and transmitter system parameter states; S p This is a set of states for each receiver, including receiver received power, number of receiving antennas, and receiver system parameter states; S n A is a set of channel states, including ambient noise, channel fading, and signal-to-noise ratio; d Let A be the set of actions for a relay node to travel a straight distance Δd. l Let A be the set of actions for a relay node to turn left and travel a distance Δd. r Let A be the set of actions for a relay node to turn right and travel a distance Δd. s This is the set of actions that stop the relay node from moving.
[0031] Specifically, in S1, firstly, the end-to-end approximate average signal-to-noise ratio of the information and energy co-transmission of the MIMO system is calculated; then, with the optimization objective of maximizing the approximate average signal-to-noise ratio between the source node and the relay node and between the relay node and the destination node, a triplet model is established using deep reinforcement learning to solve for W (0 < W ≤ M) optimal relay node positions, and W optimal or approximate optimal relay nodes are selected based on these W positions.
[0032] To quantify the information and energy co-transmission performance of MIMO systems, this invention uses the end-to-end approximate average signal-to-noise ratio (SNR) of the MIMO system as the criterion for judging system performance. In the downlink, from the source node to the relay node r... m Then go to the destination node u k The end-to-end approximate average signal-to-noise ratio expression for:
[0033]
[0034] Where m = 1, 2, ..., M represents the m-th link from the source node to the destination node, and M represents the total number of relay nodes; k = 1, 2, ..., K represents the k-th destination node, and K represents the total number of destination nodes; and Let represent the signal-to-noise ratio (SNR) from the source node to the relay node and from the relay node to the destination node, respectively, as expressed below:
[0035]
[0036] In the formula, P BS and These represent the transmit power of the source node and the relay node, respectively, calculated based on the historical average transmit power of the feasible link from the source node to the destination node. and These are the average channel fading coefficients of feasible links between the source node and the m-th relay node, and between the m-th relay node and the destination node, respectively. and These are the historical average noise power of the feasible link received by the relay node and the destination node, respectively.
[0037] During relay node deployment, based on the definition of end-to-end approximate average signal-to-noise ratio (SNR), by identifying the W relay node locations that maximize the approximate average SNR between the source node and the relay node, and between the relay node and the destination node, and then selecting W optimal or near-optimal relay nodes from these W locations, the performance of information and energy co-transmission services for MIMO systems can be maximized. Therefore, the optimization objective can be defined as maximizing the sum of the approximate average SNR between the source node and the relay node, and between the relay node and the destination node.
[0038]
[0039] Furthermore, this invention employs the framework of Deep Q-Network (DQN), combining deep learning neural networks with reinforcement learning Q-learning. It finds the optimal relay node position by maximizing cumulative reward, while utilizing the DoubleDQN algorithm to reduce the tendency for overestimation of Q-values in traditional DQN algorithms. This method requires defining the basic triples: state, action, and reward, specifically as follows:
[0040] State: The state S at time t t It is a set of transmitter states, receiver states, and channel states in the information and energy coordinated transmission of MIMO systems:
[0041] S t ={S v S r S n}
[0042] S v This is a set of states for each transmitter, including transmitter power, number of transmit antennas, and transmitter system parameter states, reflecting the performance of the transmitter; S p This is a set of states for each receiver, including receiver power, number of receiving antennas, and receiver system parameter states, reflecting the receiver's performance; S n It is a set of channel states, including ambient noise, channel fading, and signal-to-noise ratio, which can reflect the environmental conditions and overall performance.
[0043] Action: Action A at time t t It is a set of actions for node mapping and resource allocation:
[0044] A t ={A d A l A r A s}
[0045] In the formula, A d For the set of actions where the relay node moves a straight distance Δd, adjust the relay node's position forward; A l Let A be the set of actions for a relay node to turn left and travel a distance Δd. Adjust the relay node's position to the left. r Let A be the set of actions for a relay node to turn right and travel a distance Δd. Adjust the relay node's position to the right. sThe set of actions for stopping a relay node is defined, and the relay node position is fixed. By continuously training the neural network, the probability of selecting a valid action can be increased, ultimately obtaining the optimal relay node position that maximizes the sum of the approximate average signal-to-noise ratios between the source node and the relay node, and between the relay node and the destination node.
[0046] Reward: Reward R t Depends on state S t And Action A t This should be related to the previous optimization objective. The optimization objective is to maximize the approximate average signal-to-noise ratio between the source node and the relay node, and between the relay node and the destination node. The Double DQN algorithm finds the optimal strategy by maximizing the cumulative reward, so the reward r at time t can be considered as... t Defined as the sum of the approximate average signal-to-noise ratios between the source node and the relay node, and between the relay node and the destination node at time t:
[0047]
[0048] In the formula, L is a positive number. If an invalid action is selected, such as the destination node not receiving any signal, the reward for that action can be set to -200, i.e., L = 200.
[0049] The Double DQN algorithm consists of two similar convolutional neural networks and an empirical replay pool. Its principle is to use the estimated value Q(s) of the neural network. t a t ,ω) to approximate the optimal action value function Q * (s t a t Based on the current environment, the agent continuously tries various actions to obtain rewards, thereby changing the environment to the next state and collecting multiple sets of [s] t a t r t s t+1 The vectors are placed into the experience replay pool. Compared to directly obtaining continuous training data sets, randomly selecting a batch of vector samples from the experience replay pool can reduce the correlation between data sets. The convolutional neural network is trained using the randomly selected vector sets, and the parameters of the neural network are continuously updated through training to gradually improve the fitting accuracy. When the training reaches a certain number of iterations, the neural network can fit the optimal action-value function well, thus obtaining the optimal deployment strategy.
[0050] When a service request arrives, the initial state of the MIMO system is first determined by exploring the information and energy co-transmission. Let's assume the state at time t is s. t Using an ε-greedy strategy: randomly select an action a with probability ε. t s with a probability of 1-ε tInput the main neural network and select action 'a' based on the estimate from the action value function. t Execute a t Receive the corresponding reward r t And cause the environment to change to the next state s t+1 A set of vectors [s] is obtained. t a t r t s t+1 Store it in the experience pool. Then, from state s... t+1 Begin the exploration and proceed to the next action, a. t+1 The deployment cycle ends when all relay nodes are deployed, and then the next deployment cycle starts from the beginning, repeating this process to enrich the experience replay pool.
[0051] Once the experience replay pool contains a sufficient number of vectors, a batch of vectors can be randomly drawn from the experience replay pool. Taking one set of vectors [s] as an example... t a t r t s t+1 For example, let's take s as an example. t and a t Inputting the main neural network yields the Q estimate Q(s). t a t ω), where ω represents the weights of the main neural network. Find the action corresponding to the maximum Q-value in the current Q-network:
[0052] a′ t+1 =argmax a Q(s t+1 a t ,ω)
[0053] a′ t+1 and s t+1 Input the target Q network, plus r t Obtain the target Q value:
[0054]
[0055] ω - The weights of the target Q-network are represented by the value of the target Q-network. The loss function is obtained by subtracting the Q-value of the main network from the Q-value of the target network.
[0056]
[0057] The weights ω of the main network are updated using gradient descent with the loss function Loss. The weights ω of the target Q-network are... - The updates depend on the training results of the main network. To reduce the impact of data fluctuations on model stability, every time interval T, the weights ω of the main network are copied to the weights ω of the target network. -This process keeps the target network constant for a period of time, reducing model volatility. This process is repeated iteratively until the loss function converges to a sufficiently small range. At this point, the main neural network has converged, and the optimal strategy can be obtained by deploying the system based on the Q-evaluation.
[0058] like Figure 3 As shown, the optimal relay node location is determined using the Double DQN algorithm, and the optimal relay node is determined based on this location. This process includes the following steps:
[0059] 1. Initialize the experience replay pool, randomly initialize the master Q-network, and copy the parameters to the target Q-network, i.e., ω. - =ω.
[0060] 2. Initialize the state set S t ={S v S r S n}, and let t = 0, to obtain state s t ;
[0061] 3. Determine the set of optional actions A t Randomly select an action a with probability ε. t The optimal action a is selected with a probability of 1-ε based on the estimate of the main neural network Q. t =argmax a Q(s t a t ,ω).
[0062] 4. Perform action a t Receive reward r t , reach the next state s t+1 , will vector [s t a t r t s t+1 Add the data to the experience replay pool. If the experience replay pool is full, proceed to step 5; if the experience replay pool is not full, but all relay nodes are in a stopped state, proceed to step 2; if the experience replay pool is not full, and there are relay nodes that are not in a stopped state, then t = t + 1, and proceed to step 3.
[0063] 5. Randomly select a batch of vectors from the experience replay pool.
[0064] 6. Extract vectors [s] one by one from it. t a t r t s t+1The vector is then input into the main network, and the action corresponding to the maximum Q-value is found in the main network. This action is then combined with the aforementioned vector and input into the target network to obtain the target network Q-value under the new model. The loss function is then calculated using the obtained main network Q-value and the target network Q-value.
[0065]
[0066] The parameter ω is updated using the loss function via gradient descent.
[0067] 7. Repeat steps 5-6 and copy the weights of the main network to the target network ω every time interval T. - =ω, until the network converges or the preset number of iterations is reached, thus obtaining the optimal relay node position.
[0068] 8. Obtain the optimal relay node location in the intelligent information and energy coordinated transmission of the MIMO system through the main network. If the location Q(s) t a t If ω is negative, the request is rejected.
[0069] 9. Based on the W optimal relay node locations obtained in step 8, query whether there is a node at each location in the system. If there is, select the node as the optimal relay node; if there is no node, select the node closest to the location as the approximate optimal relay node.
[0070] The method provided by this invention, in step S1, employs the Double DQN algorithm from deep reinforcement learning. It obtains the optimal relay node position by maximizing the approximate average signal-to-noise ratio between the source node and the relay node, and between the relay node and the destination node, and selects the optimal or near-optimal node based on this position. The Double DQN algorithm reduces the likelihood of overestimating the Q-value in traditional DQN algorithms. Initially, W existing relay node positions in the network are randomly selected. For different user terminals, the Double DQN algorithm continuously updates the relay node positions until the constraints are met, finally obtaining W optimal relay node positions. It is understood that the specific value of W can be set according to the actual situation.
[0071] S2, determine the optimal transmit power of the source node and the optimal relay forwarding power of the optimal relay node.
[0072] Based on the selected optimal or near-optimal relay node, the optimal transmit power is calculated under the constraint of maximizing system energy efficiency. The solution considers cases where the signal-to-noise power ratio is high, and the calculation formula is simplified based on logarithmic function theory. The optimal transmit power includes the optimal transmit power of the source node and the relay forwarding power of the optimal relay node. Under the constraint of maximizing system energy efficiency, the optimal relay forwarding power is first calculated, and the relay channel energy efficiency function is expressed as the relay forwarding power p. m The function is solved considering the case where the ratio of signal to noise power is high. Based on the theory of logarithmic functions, η is simplified. EE (p m ),as follows:
[0073]
[0074] Where, p c =Kp1+2p syn +Mp2, p1, p2, p syn These represent the power consumed by the transmitter's RF circuit, the power consumed by the receiver's RF circuit, and the power consumed for system synchronization, respectively; K and M represent the total number of destination nodes and the total number of relay nodes, respectively; β k Let be the large-scale fading coefficient of the k-th destination node; B represents the system bandwidth. The optimal relay forwarding power can be calculated by maximizing energy efficiency, that is, by maximizing the energy efficiency function η of the forwarding channel. EE (p m The relay forwarding power p at its maximum m This represents the optimal relay forwarding power for the m-th optimal relay node.
[0075] Since the energy efficiency function from the source node to the relay node is similar to the above formula, the optimal source node transmit power can be obtained by performing constraint calculations again. When solving, the case of a high signal-to-noise power ratio is considered. Based on logarithmic function theory, η is simplified. EE (p n ).
[0076] Specifically:
[0077]
[0078]
[0079] Where, p′ c =Mp′1+2p′ syn +Np′2,p′1、p′2、p′ syn These represent the power consumed by the RF circuitry at the source transmitter, the power consumed by the RF circuitry at the relay receiver, and the power consumed for system synchronization, respectively; N and M represent the total number of MIMO transmitter antennas and the total number of relay nodes, respectively; βm Let be the large-scale fading coefficient of the m-th relay node; B represents the system bandwidth. The optimal source transmit power can be calculated by maximizing energy efficiency, that is, by maximizing the source transmit channel energy efficiency function η. EE (p n The source emission power p at its maximum n The optimal transmit power for the nth MIMO transmitter antenna.
[0080] S3 determines the beamforming vector with the goal of maximizing the spectral efficiency of the target node.
[0081] Specifically, with the goal of maximizing the spectral efficiency of the target node, intelligent hybrid beamforming is used to compensate for signal propagation loss in free space, including:
[0082] Training Phase: The channel matrices from the source node to the relay node and from the relay node to the destination node are used as inputs to the neural network. The beamformer is optimized by minimizing a defined loss function using stochastic gradient descent, aiming to bring the spectral efficiency of the destination node as close as possible to its optimal performance. The loss function is defined using spectral efficiency, ensuring that decreasing the loss function results in an increase in spectral efficiency, as shown below:
[0083]
[0084] Where N represents the total number of training data samples, (λ n H n F RF,n, W RF,n Let represent the signal-to-noise ratio, channel matrix, and analog beamforming matrices of the transmitter and receiver, respectively, for the nth sample. N represents s Self-information of each data stream, N s The number of data streams sent by the MIMO downlink system.
[0085] Testing phase: The trained model replaces the beamformer, taking the channel matrix vector as input and outputting the beamforming vector. Using the output beamforming vector, the orientation and state of multiple antennas are adjusted to maximize the spectral efficiency of the target node.
[0086] This invention provides an intelligent information and energy collaborative transmission deployment device for MIMO systems, comprising:
[0087] S1, the first processing module, is used to determine the optimal relay node location using the Double DQN algorithm with the goal of maximizing the approximate average signal-to-noise ratio from the source node to the relay node and from the relay node to the destination node in the MIMO system, and to determine the optimal relay node based on the location.
[0088] Among them, St ={S v S r S n}、A t ={A d A l A r A s The approximate average signal-to-noise ratios from the source node to the relay node and from the relay node to the destination node at time t are respectively used as the state set, action set, and reward value at time t; S v This is a set of states for each transmitter, including transmitter power, number of transmit antennas, and transmitter system parameter states; S p This is a set of states for each receiver, including receiver received power, number of receiving antennas, and receiver system parameter states; S n A is a set of channel states, including ambient noise, channel fading, and signal-to-noise ratio; d Let A be the set of actions for a relay node to travel a straight distance Δd. l Let A be the set of actions for a relay node to turn left and travel a distance Δd. r Let A be the set of actions for a relay node to turn right and travel a distance Δd. s This is the set of actions that stop the relay node from moving. For the source node to the relay node r m r m to destination node u k The approximate average signal-to-noise ratio, U K R is the set of all destination nodes. M For the set of all relay nodes;
[0089] S2, the second processing module, is used to determine the optimal transmit power of the source node and the optimal forwarding power of the optimal relay node;
[0090] S3, the third processing module, is used to determine the beamforming vector with the goal of maximizing the spectral efficiency of the target node.
[0091] This invention provides an intelligent information and energy collaborative transmission deployment system for MIMO systems, comprising: a computer-readable storage medium and a processor;
[0092] The computer-readable storage medium is used to store executable instructions;
[0093] The processor is configured to read executable instructions stored in the computer-readable storage medium and execute the method as described in any of the above embodiments.
[0094] This invention provides a computer-readable storage medium storing computer instructions that cause a processor to perform the method described in any of the above embodiments.
[0095] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for deploying intelligent information and energy coordinated transmission in MIMO systems, characterized in that, include: S1, with the goal of maximizing the approximate average signal-to-noise ratio from the source node to the relay node and from the relay node to the destination node in the MIMO system, uses the Double DQN algorithm to determine the optimal relay node location, and determines the optimal relay node based on this location. Among them, with , and t The approximate average signal-to-noise ratio from the source node to the relay node and from the relay node to the destination node at time t are respectively used as the state set, action set, and reward value at time t; This is a set of states for each transmitter, including transmitter power, number of transmitting antennas, and transmitter system parameter states; This is a set of states for each receiver, including receiver receiving power, number of receiving antennas, and receiver system parameter states; This is a set of channel states, including ambient noise, channel fading, and signal-to-noise ratio. For relay nodes to proceed straight The set of actions related to distance. Turn left at the relay node The set of actions related to distance. Turn right at the relay node The set of actions related to distance. This is the set of actions that stop the relay node from moving. S2, determine the optimal transmit power of the source node and the optimal forwarding power of the optimal relay node; S3, with the goal of maximizing the spectral efficiency of the target node, determines the beamforming vector; In step S1, the approximate average signal-to-noise ratio from the source node to the relay node and from the relay node to the destination node is: in, From source node to relay node , to the destination node Approximate average signal-to-noise ratio; Represents the first step from the source node to the destination node. One feasible link, This represents the total number of relay nodes; Representing the first One destination node, This represents the total number of destination nodes; and These represent the signal-to-noise ratios from the source node to the relay node and from the relay node to the destination node, respectively. , and These represent the transmit power of the source node and the relay node, respectively. and From the source node to the 1st The relay node and the first The average channel fading coefficient of a feasible link between a relay node and a destination node; and These are the historical average noise power of the feasible link received by the relay node and the destination node, respectively. In step S2, the forwarding channel energy efficiency function is... Maximum relay forwarding power As the first The optimal relay forwarding power of the optimal relay node; in, , , , , These represent the power consumed by the radio frequency circuit at the relay transmitter, the power consumed by the radio frequency circuit at the receiver, and the power consumed for system synchronization, respectively. and These represent the total number of destination nodes and the total number of relay nodes, respectively. For the first Large-scale fading coefficient of each target node; Indicates system bandwidth; Source transmit channel energy efficiency function Maximum source transmit power As the first The optimal transmit power of each MIMO transmitter antenna; in, , , , , These represent the power consumed by the radio frequency circuit at the source transmitter, the power consumed by the radio frequency circuit at the relay receiver, and the power consumed for system synchronization, respectively. and These represent the total number of MIMO transmitter antennas and the total number of relay nodes, respectively. For the first Large-scale fading coefficient of each relay node; Indicates system bandwidth; In step S3, the channel matrix from the source node to the relay node and the channel matrix from the relay node to the destination node are input into the pre-trained neural network to obtain the beamforming vector.
2. The method as described in claim 1, characterized in that, The loss function of the pre-trained neural network during the training phase is: ; in, This represents the total number of training data samples. They represent the first The signal-to-noise ratio, channel matrix, analog beamforming matrix at the transmitter, and analog beamforming matrix at the receiver are for each sample. express Self-information of a data stream The number of data streams sent by the MIMO downlink system.
3. The method as described in claim 1, characterized in that, In step S1, if a node exists at the optimal relay node location, it is taken as the optimal relay node; if it does not exist, the node closest to that location is taken as the optimal relay node.
4. An intelligent information and energy collaborative transmission deployment device for MIMO systems, characterized in that, include: S1, the first processing module, is used to determine the optimal relay node location using the Double DQN algorithm with the goal of maximizing the approximate average signal-to-noise ratio from the source node to the relay node and from the relay node to the destination node in the MIMO system, and to determine the optimal relay node based on the location. Among them, with , and t The approximate average signal-to-noise ratio from the source node to the relay node and from the relay node to the destination node at time t are respectively used as the state set, action set, and reward value at time t; This is a set of states for each transmitter, including transmitter power, number of transmitting antennas, and transmitter system parameter states; This is a set of states for each receiver, including receiver receiving power, number of receiving antennas, and receiver system parameter states; This is a set of channel states, including ambient noise, channel fading, and signal-to-noise ratio. For relay nodes to proceed straight The set of actions related to distance. Turn left at the relay node The set of actions related to distance. Turn right at the relay node The set of actions related to distance. This is the set of actions that stop the relay node from moving. S2, the second processing module, is used to determine the optimal transmit power of the source node and the optimal forwarding power of the optimal relay node; S3, the third processing module, is used to determine the beamforming vector with the goal of maximizing the spectral efficiency of the target node; The approximate average signal-to-noise ratio from the source node to the relay node and from the relay node to the destination node is: in, From source node to relay node , to the destination node Approximate average signal-to-noise ratio; Represents the first step from the source node to the destination node. One feasible link, This represents the total number of relay nodes; Representing the first One destination node, This represents the total number of destination nodes; and These represent the signal-to-noise ratios from the source node to the relay node and from the relay node to the destination node, respectively. , and These represent the transmit power of the source node and the relay node, respectively. and From the source node to the 1st The relay node and the first The average channel fading coefficient of a feasible link between a relay node and a destination node; and These are the historical average noise power of the feasible link received by the relay node and the destination node, respectively. forwarding channel energy efficiency function Maximum relay forwarding power As the first The optimal relay forwarding power of the optimal relay node; in, , , , , These represent the power consumed by the radio frequency circuit at the relay transmitter, the power consumed by the radio frequency circuit at the receiver, and the power consumed for system synchronization, respectively. and These represent the total number of destination nodes and the total number of relay nodes, respectively. For the first Large-scale fading coefficient of each target node; Indicates system bandwidth; Source transmit channel energy efficiency function Maximum source transmit power As the first The optimal transmit power of each MIMO transmitter antenna; in, , , , , These represent the power consumed by the radio frequency circuit at the source transmitter, the power consumed by the radio frequency circuit at the relay receiver, and the power consumed for system synchronization, respectively. and These represent the total number of MIMO transmitter antennas and the total number of relay nodes, respectively. For the first Large-scale fading coefficient of each relay node; Indicates system bandwidth; The channel matrix from the source node to the relay node and the channel matrix from the relay node to the destination node are input into a pre-trained neural network to obtain the beamforming vector.
5. An intelligent information and energy collaborative transmission deployment system for MIMO systems, characterized in that, include: Computer-readable storage media and processors; The computer-readable storage medium is used to store executable instructions; The processor is configured to read executable instructions stored in the computer-readable storage medium and execute the method as described in any one of claims 1-3.
6. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to perform the method as described in any one of claims 1-3.