RIS-aided mobile industrial robot wireless power transfer optimization method based on td3 algorithm
By optimizing the ratio of RIS reflection energy and signal through the TD3 algorithm, the signal blocking and battery capacity limitations in wireless industrial robot communication are solved, thereby improving communication quality and battery life.
Patent Information
- Application Number
- CN202510381181.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-03-28
AI Technical Summary
Wireless industrial robot communication faces problems such as signal congestion and battery capacity limitations, resulting in poor communication quality and short battery life.
A wireless energy-carrying communication method for RIS-assisted mobile industrial robots based on the TD3 algorithm is adopted. By modeling the channel and path loss, the spatial and temporal ratio of the reflected energy and signal of the RIS is optimized using the TD3 reinforcement learning model to maximize energy and communication quality.
While meeting the minimum communication quality threshold, the communication quality and battery life of wireless industrial robots have been improved.
Smart Images

Figure CN120302316B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the field of communication, and particularly relates to a RIS-assisted mobile industrial robot wireless energy-carrying communication optimization method based on a TD3 algorithm. BACKGROUND
[0002] Wireless communication technology and robot technology bring many conveniences to industry 4.0. Compared with wired communication, industrial wireless communication is easy to deploy and maintain, and can reduce costs by 90%; and can realize the free movement of field devices, flexible network structure and the diversity of applications. At the same time, wireless communication-assisted mobile industrial robots can be clustered without being limited by the workspace, realizing diversified manufacturing and adapting to different environments, which helps to reduce costs.
[0003] However, the application of wireless industrial robots faces two major problems. First, key tasks of industrial production require deterministic communication. Because of mobile objects (such as industrial robots, trucks, mechanical arms, etc.) and densely deployed metal machines, wireless signals are more susceptible to blockage. In addition, battery capacity also limits the performance of industrial robots, such as working range, service time, payload, etc.
[0004] Existing work has not solved the above two major problems. Therefore, how to change the position of the mobile industrial robot, the dynamic occlusion situation, meet the minimum threshold of communication quality, deploy RIS to reflect energy, the spatial proportion of signals, and the time proportion of simply reflecting energy and reflecting energy and signals at the same time, and realize the maximization of transmission energy and communication quality through the intelligent decision of the TD3 reinforcement learning model, has become a very important problem. SUMMARY
[0005] The present application aims at the deficiencies of the prior art, and provides a RIS-assisted mobile industrial robot wireless energy-carrying communication optimization method based on a TD3 algorithm.
[0006] The purpose of the present application is achieved by the following technical scheme: a RIS-assisted mobile industrial robot wireless energy-carrying communication optimization method based on a TD3 algorithm, the optimization method is based on a RIS-assisted mobile industrial robot wireless energy-carrying communication system; the communication system includes an access point, a RIS and industrial robots; the access point has only one antenna with D arrays for beamforming, which can simultaneously emit energy, signal electromagnetic waves; the industrial robots are K, denoted as respectively installed with a single array antenna; the RIS has only one, which receives energy, signal electromagnetic waves from the access point and reflects them to the K industrial robots; the RIS-assisted mobile industrial robot wireless energy-carrying communication optimization method based on the TD3 algorithm includes:
[0007] (1) model the channel between the access point and the RIS, and the channel between the RIS and the K robots;
[0008] (2) model the channel path loss between the access point and the RIS, and the channel path loss between the RIS and the K robots;
[0009] (3) establish an optimization problem for energy transmission and information transmission at different time slots t according to the channel path loss;
[0010] (4) establish a deep reinforcement learning model according to the channel environment, the variables to be optimized, and the objectives to be optimized;
[0011] (5) optimize the reinforcement learning model using the TD3 algorithm, and obtain the solution to the optimization problem according to the optimized reinforcement learning model to obtain the time proportion and spatial proportion of deploying the RIS for reflecting energy and signals at different time slots t.
[0012] Further, the modeling of the channel between the access point and the RIS, and the channel between the RIS and the K robots includes:
[0013] (1) establish a three-dimensional Cartesian coordinate system for all communication nodes;
[0014] (2) the RIS is equipped with m rows and n columns, a total of reflector units, and the reflector unit array is denoted as the i-th row j-th reflector unit is denoted as where i∈{1,2,…,m},j∈{1,2,…,n}, the amplitude and phase of each reflector unit can be adjusted;
[0015] (3) divide the entire time period into T equal time slots, denoted as where each time slot t is divided into two stages, an energy transmission stage and an information transmission stage; in the energy transmission stage, all RIS units are used for energy transmission, and the time proportion in the time slot t is τ(t), 0≤τ(t)≤1; in the information transmission stage, only RIS units are used for information transmission for industrial robot k, denoted as and the other part of the RIS units are used for energy transmission; the time proportion of the information transmission stage in the time slot t is 1-τ(t); the proportion of the number of RIS units used for information transmission in the information transmission stage to the total number of RIS units is denoted as
[0016] (4) the beamforming of the access point adopts a linear precoding strategy, and the transmitted signal is represented as:
[0017]
[0018] in, and These are electromagnetic waves for energy transmission and electromagnetic waves for information transmission, respectively. and These are the precoded vectors used by the access point to transmit energy and information, respectively. Represents a complex vector in a D×1 space; and The original signals used by the access point for transmitting energy and information are respectively, both following a cyclic symmetric complex Gaussian distribution with mean 0 and variance 1; the total power p used by the access point for transmitting energy is expressed as:
[0019]
[0020] in, represents the conjugate transpose, and ||·|| represents the 2-norm of the vector;
[0021] (5) In each time slot t, the industrial robot The energy received E k (t) is represented as:
[0022]
[0023] Among them, Z k From the access point to the RIS unit The baseband equivalent channel; From RIS unit The baseband equivalent channel to industrial robot k; It is a RIS unit The diagonal matrix of the reflection coefficient matrix, where θ l and β l These represent the phase and amplitude of the RIS unit l, respectively; From the access point to part of the RIS unit The baseband equivalent channel; From a portion of the RIS unit The baseband equivalent channel to industrial robot k; It is a part of the RIS unit The diagonal matrix of the reflection coefficient matrix; η∈(0,1) is the energy transfer efficiency; ω i,j;k (t) is an indicator function that indicates the RIS unit. Whether it is used for transmitting information is indicated as follows:
[0024]
[0025] (6) During the energy transfer phase in each time slot t, where the proportion is τ(t) and 0≤τ(t)≤1, all RIS units They are evenly distributed to each industrial robot k;
[0026] (7) In each time slot t, the industrial robot The received information y k (t) is represented as:
[0027]
[0028] Among them, v k The noise level of each industrial robot is represented by the noise power formula. Additive white Gaussian noise;
[0029] (8) Considering both line-of-sight and non-line-of-sight channels, the channel model from RIS to industrial robot is modeled as following a Rayleigh distribution.
[0030] Furthermore, the modeling of the channel between the access point and the RIS, and the channel path loss between the RIS and the K robots, includes:
[0031] (1) The channel between the access point and the RIS is mainly a line-of-sight (LoS) channel. Channel path loss between The model is as follows:
[0032]
[0033] Among them, (x i,j ,y i,j ,z i,j ) represents the RIS unit in the i-th row and j-th column. The location in the three-dimensional Cartesian coordinate system; the access point is at the origin of the three-dimensional Cartesian coordinate system; α represents the distance from the access point to the RIS unit. The channel path loss exponent; κ represents the channel path loss at the reference distance d′;
[0034] (2) The access point is assumed to be a point mass in three-dimensional space;
[0035] (3) The channel between the RIS and the industrial robot includes a line-of-sight channel and a non-line-of-sight (NLoS) channel. Channel path loss L RI The model is as follows:
[0036]
[0037] in, These represent the probabilities of the LoS channel and the NLoS channel between the RIS and the industrial robot, respectively. This represents the channel path loss index from the RIS to the industrial robot; Indicates the additional path loss factor under NLoS channel; d k (t) represents the distance from RIS to each industrial robot k in each time slot t;
[0038] (4) In the channel between the RIS and the industrial robot, the d in each time slot t k (t) is considered a constant;
[0039] (5) The LoS channel between the RIS and the industrial robot is modeled as follows:
[0040]
[0041] Wherein, variable X represents the number of times the channel between the RIS and the target industrial robot is blocked by other mobile industrial robots per unit time, and X follows a Poisson distribution with parameter λ; the probability distribution function of X is expressed as:
[0042]
[0043] Where λ is proportional to unit time, denoted as:
[0044]
[0045] Where c, c>0 is a scaling factor used to represent the severity of the blockage;
[0046] (6) The NLoS channel between the RIS and the industrial robot is modeled as follows:
[0047]
[0048] (7) Considering radio frequency interference, the signal-to-interference-plus-noise ratio (SINR) is used to represent the industrial robot. The quality of the received signal is defined as:
[0049]
[0050] Among them, SINR k (t) represents the signal-to-interference-plus-noise ratio (SIR) of the k-th industrial robot in time slot t.
[0051] Furthermore, the optimization problem for energy transmission and signal transmission under different time slots t, based on channel path loss, includes:
[0052] The system objective is to optimize the spatial ratio ξ of reflected energy and signal by using the RIS unit. k(t), and the time ratio τ(t) / (1-τ(t)) of the reflected energy alone and the reflected energy to the signal simultaneously, at the minimum threshold SINR for meeting communication quality. min To maximize transmission energy and communication quality, the optimization problem can be formulated as follows:
[0053]
[0054] in, Let be the reward function for the k-th industrial robot in time slot t, δ be the energy and signal balance coefficients used to characterize the importance of communication quality in the overall optimization objective, and C1 be the minimum threshold for communication quality, i.e., the minimum signal-to-interference-plus-noise ratio constraint SINR. min C2 represents the finite-time constraint; C3 represents the maximum energy transfer power constraint p. max C4 represents the RIS ratio constraint.
[0055] Furthermore, the step of establishing a deep reinforcement learning model based on the channel environment, the variable to be optimized, and the objective to be optimized includes:
[0056] (1) The state space of the deep reinforcement learning model includes: the transmitted signal G of the access point; the distance d between RIS and the k-th industrial robot. k (t); The baseband equivalent channel h of RIS and the k-th industrial robot k (t);
[0057] (2) The action space of the deep reinforcement learning model includes: the time proportion and the spatial proportion ξ under time slot t. k (t);
[0058] (3) The reward function of the deep reinforcement learning model is modeled as: the optimization objective of claim 4.
[0059] (4) The optimization objective The received dryness ratio must not be lower than the minimum threshold SINR. min Due to the constraints, the reward value of the k-th industrial robot in time slot t is... The model is as follows:
[0060]
[0061] Furthermore, in step (5), the reinforcement learning model is optimized using the TD3 algorithm, and the solution to the optimization problem is obtained based on the optimized reinforcement learning model. This yields the time ratio τ(t) and spatial ratio ξ of RIS deployment for reflecting energy and signal under different time slots t. k (t), including:
[0062] (1) Initialize the Actor policy network Critic Review Network And Critic Review Network 2
[0063] and the corresponding Actor target network and Critic target network and Critic target network Where s represents the state, a represents the action, and w represents the network parameter vector;
[0064] (2) The number of training rounds (episode) is initialized to 0;
[0065] (3) The time slot t in the training episode is initialized to 0;
[0066] (4) The Actor policy network outputs action A(t) and obtains a reward value based on the input state S(t). Then, the system transitions to the next state S(t+1) to obtain the training dataset. And store it in an experience replay repository of size B. middle;
[0067] (5) Randomly sample B from the experience replay repository. m A dataset consisting of samples is sent to the Actor policy network, Critic review network 1, Critic review network 2, Actor target network, Critic target network 1, and Critic target network 2; where the i-th sample is denoted as
[0068] (6) The objective function is calculated from the minimum Q-value function of Critic objective network 1 and Critic objective network 2:
[0069]
[0070] Among them, y i γ is the objective function value of the i-th sample; γ is the penalty factor, representing the importance of the expected Q-value function reward in the future. For the bounded smoothed target action, by giving action A i+1 Add bounded noise To obtain;
[0071] (7) Critic review network 1 and Critic review network 2 are obtained by minimizing the mean square error between the above objective function and the expected Q value:
[0072]
[0073] Where j = 1, 2 correspond to Critic review network 1 and Critic review network 2, respectively;
[0074] (8) When time interval t passes through T u Next, update the Actor policy network; Actor policy network parameters. The update is obtained by maximizing the expected Q value, and the gradient of the expected Q value is expressed as:
[0075]
[0076] (9) When time interval t passes through T u Next, update the target network; based on the Actor policy network parameters... Critic Review Network 1 Parameter And Critic review network 2 parameters Smoothly update the Actor target network parameters respectively Critic target network 1 parameter And Critic target network 2 parameters
[0077]
[0078] Where ψ∈[0,1] are the smooth update hyperparameters;
[0079] (10) Determine whether the time slot t < T, where T is the total number of time slots; if yes, then t = t + 1 and return to step (4); otherwise, complete the RIS policy generation for one round and continue training for the next round.
[0080] (11) Determine whether the number of rounds episode < EPISODE, where EPISODE is the total number of rounds; if yes, then episode = episode + 1, and return to step (3); otherwise, the optimization ends and the optimized reinforcement learning model is obtained.
[0081] Furthermore, the target action after bounded smoothing By giving action A i+1 Add bounded noise To obtain:
[0082]
[0083] Among them, A i+1 The output of the Actor target network, To represent a bounded function, Restricted to the lower bound α L and upper bound α HBetween, and then output to That is: if but if but if but Represented as:
[0084]
[0085] Among them, bounded noise Defined as c and -c are bounded functions clip pairs, respectively. The upper and lower bounds; This indicates that the expression follows a pattern with a mean of 0 and a variance of . It follows a normal distribution.
[0086] The present invention also provides an electronic device, including a memory and a processor, wherein the memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the above-described RIS-assisted mobile industrial robot wireless power-carrying communication optimization method based on the TD3 algorithm.
[0087] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the above-described RIS-assisted mobile industrial robot wireless power-carrying communication optimization method based on the TD3 algorithm.
[0088] The present invention also provides a computer program product, including a computer program / instruction, characterized in that, when the computer program / instruction is executed by a processor, it implements the above-mentioned RIS-assisted mobile industrial robot wireless power-carrying communication optimization method based on the TD3 algorithm.
[0089] The beneficial effects of this invention are as follows: This invention provides a RIS-assisted wireless energy-carrying communication optimization method for mobile industrial robots based on the TD3 algorithm. According to the positional changes and dynamic occlusion of the mobile industrial robot, while meeting the minimum communication quality threshold, it maximizes transmission energy and communication quality by deploying RIS for the spatial ratio of reflected energy and signal, and the temporal ratio of purely reflected energy and simultaneously reflected energy and signal, through intelligent decision-making using the TD3 reinforcement learning model. This invention also solves the problems of poor wireless communication quality and short battery life for industrial robots. Attached Figure Description
[0090] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0091] Figure 1 This is a diagram of the wireless power-carrying communication model for a RIS-assisted mobile industrial robot provided in an embodiment of the present invention.
[0092] Figure 2 This is a diagram of the TD3 algorithm framework provided in an embodiment of the present invention;
[0093] Figure 3 This is the reward graph of the TD3 algorithm provided in this embodiment of the invention under the number of training steps;
[0094] Figure 4 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0095] To better understand the technical solution of this application, the embodiments of this application will be described in detail below with reference to the accompanying drawings.
[0096] It should be understood that the described embodiments are merely some, not all, of the embodiments in this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.
[0097] This invention provides a RIS-assisted wireless energy-carrying communication optimization method for mobile industrial robots based on the TD3 algorithm. According to the positional changes and dynamic occlusion of the mobile industrial robot, while meeting the minimum communication quality threshold, it maximizes transmission energy and communication quality by deploying RIS for the spatial ratio of reflected energy and signal, and the temporal ratio of purely reflected energy versus simultaneously reflected energy and signal, through intelligent decision-making using a TD3 reinforcement learning model. This invention also solves the problems of poor wireless communication quality and short battery life for industrial robots.
[0098] This optimization method is based on, for example Figure 1 The diagram illustrates a RIS-assisted wireless energy-carrying communication system for mobile industrial robots. The system includes an access point, a RIS, and industrial robots. The access point has only one antenna with D arrays for beamforming, capable of simultaneously transmitting energy and signal electromagnetic waves. There are K industrial robots, denoted as K. Each is equipped with a single array of antennas; there is only one RIS, which receives energy and electromagnetic waves from the access point and reflects them to K industrial robots; the RIS-assisted mobile industrial robot wireless energy-carrying communication optimization method based on the TD3 algorithm includes:
[0099] Step 1: Model the channels between the access point and the RIS, and between the RIS and the K robots; specifically:
[0100] (1) Establish a three-dimensional Cartesian coordinate system for all communication nodes;
[0101] (2) RIS is equipped with m rows and n columns, total There are 1 reflective units, and the array of reflective units is denoted as . The reflection unit in the i-th row and j-th column is denoted as Where i∈{1,2,…,m},j∈{1,2,…,n}, and the amplitude and phase of each reflection unit can be adjusted;
[0102] (3) Divide the entire time period into T equal-length time slots, denoted as . Each time slot t is divided into two stages: an energy transmission stage and an information transmission stage. In the energy transmission stage, all RIS units are used for energy transmission, and their time allocation within time slot t is τ(t), where 0 ≤ τ(t) ≤ 1. In the information transmission stage, for the industrial robot k, only… Each RIS unit is used for transmitting information, denoted as . Another portion of the RIS units are used for energy transmission; the time allocated to the information transmission phase in time slot t is 1-τ(t); furthermore, the ratio of the number of RIS units used for information transmission in the information transmission phase to the total number of RIS units is denoted as...
[0103] (4) The beamforming at the access point employs a linear precoding strategy, and its transmitted signal is represented as follows:
[0104]
[0105] in, and These are electromagnetic waves for energy transmission and electromagnetic waves for information transmission, respectively. and These are the precoded vectors used by the access point to transmit energy and information, respectively. Represents a complex vector in a D×1 space; and These are the original signals used by the access point to transmit energy and information, respectively, both of which follow a cyclic symmetric complex Gaussian distribution with a mean of 0 and a variance of 1, i.e.: Therefore, there is a mean. Furthermore, the total power p used by the access point for energy transmission is expressed as:
[0106]
[0107] in, represents the conjugate transpose, and ||·|| represents the 2-norm of the vector;
[0108] (5) In each time slot t, the industrial robot The energy received E k (t) is represented as:
[0109]
[0110] Among them, Z k From the access point to the RIS unit The baseband equivalent channel; From RIS unit The baseband equivalent channel to industrial robot k; It is a RIS unit The diagonal matrix of the reflection coefficient matrix, where θ l and β l These represent the phase and amplitude of the RIS unit l, respectively; similarly, From the access point to part of the RIS unit The baseband equivalent channel; From a portion of the RIS unit The baseband equivalent channel to industrial robot k; It is a part of the RIS unit The diagonal matrix of the reflection coefficient matrix; η∈(0,1) is the energy transfer efficiency; ω i,j;k (t) = 0 represents the RIS unit. Used for transmitting information, and vice versa, therefore the indicator function ω i,j;k (t) is represented as:
[0111]
[0112] (6) During the energy transfer phase in each time slot t, where the proportion is τ(t) and 0≤τ(t)≤1, all RIS units They are evenly distributed to each industrial robot k;
[0113] (7) In each time slot t, the industrial robot The received information y k (t) is represented as:
[0114]
[0115] Among them, v kThe noise level of each industrial robot is represented by the noise power formula. Additive white Gaussian noise, denoted as
[0116] (8) Considering both line-of-sight and non-line-of-sight channels, the channel model from RIS to industrial robot is modeled as following a Rayleigh distribution.
[0117] Step 2: Model the channel loss between the access point and the RIS, and between the RIS and the K robots; specifically:
[0118] (1) The channel between the access point and the RIS is mainly a line-of-sight (LoS) channel. Channel path loss between The model is as follows:
[0119]
[0120] Among them, (x i,j ,y i,j ,z i,j ) represents the RIS unit in the i-th row and j-th column. The location in the three-dimensional Cartesian coordinate system; the access point is at the origin of the three-dimensional Cartesian coordinate system; α represents the distance from the access point to the RIS unit. The channel path loss exponent; κ represents the channel path loss at a reference distance d′=1m;
[0121] (2) The access point is assumed to be a point mass in three-dimensional space;
[0122] (3) The channel between the RIS and the industrial robot includes a line-of-sight (LoS) channel and a non-line-of-sight (NLoS) channel. Channel path loss L RI The model is as follows:
[0123]
[0124] in, These represent the probabilities of the LoS channel and the NLoS channel between the RIS and the industrial robot, respectively. This represents the channel path loss index from the RIS to the industrial robot; Indicates the additional path loss factor under NLoS channel; d k (t) represents the distance from RIS to each industrial robot k in each time slot t;
[0125] (4) In the channel between the RIS and the industrial robot, the d in each time slot t k (t) is considered a constant because the time slot t period is very short, during which d k The change in (t) is very small and can be ignored;
[0126] (5) The LoS channel between the RIS and the industrial robot is modeled as follows:
[0127]
[0128] Wherein, variable X represents the number of times the channel between the RIS and the target industrial robot is blocked by other mobile industrial robots per unit time, and X follows a Poisson distribution with parameter λ; the probability distribution function of X is expressed as:
[0129]
[0130] Where λ is proportional to unit time, denoted as:
[0131]
[0132] Where c, c>0 is a scaling factor used to represent the severity of the blockage, which is usually proportional to the moving speed, number, size, etc. of the industrial robot;
[0133] (6) The NLoS channel between the RIS and the industrial robot is modeled as follows:
[0134]
[0135] (7) Considering radio frequency interference, the signal-to-interference-plus-noise ratio (SINR) is used to represent the industrial robot. The quality of the received signal is defined as:
[0136]
[0137] Among them, SINR k (t) represents the signal-to-interference-plus-noise ratio (SIR) of the k-th industrial robot in time slot t.
[0138] Step 3: Based on the channel path loss, establish an optimization problem for energy transmission and information transmission under different time slots t;
[0139] Specifically, the system objective is to optimize the spatial ratio ξ of reflected energy and signal by the RIS unit. k (t), and the time ratio τ(t) / (1-τ(t)) of the reflected energy alone and the reflected energy to the signal simultaneously, at the minimum threshold SINR for meeting communication quality. min To maximize transmission energy and communication quality, the optimization problem can be formulated as follows:
[0140]
[0141] in, Let be the reward function for the k-th industrial robot in time slot t, δ be the energy-signal balance coefficient (positive), and C1 be the minimum threshold for communication quality, i.e., the minimum signal-to-interference-plus-noise ratio (SINR) constraint. min C2 represents the finite-time constraint; C3 represents the maximum energy transfer power constraint p. max C4 represents the RIS proportionality constraint. The optimization problem described above contains the discrete variable ξ. k Given the variables τ(t) and τ(t), this optimization problem is a mixed integer nonlinear optimization problem. All possible solutions to this optimization problem suffer from state explosion, and the non-convex constraint C4 makes it difficult to solve.
[0142] Step 4: Based on the channel environment, the variables to be optimized, and the objective to be optimized, establish a deep reinforcement learning model; specifically:
[0143] (1) The state space of the deep reinforcement learning model includes: the transmitted signal G of the access point; the distance d between RIS and the k-th industrial robot. k (t); The baseband equivalent channel h of RIS and the k-th industrial robot k (t);
[0144] (2) The action space of the deep reinforcement learning model includes: the time proportion and the spatial proportion ξ under time slot t. k (t);
[0145] (3) The reward function of the deep reinforcement learning model is modeled as: the optimization objective described in step 3.
[0146] (4) The optimization objective The received dryness ratio must not be lower than the minimum threshold SINR. min Due to the constraints, the reward value of the k-th industrial robot in time slot t is... The model is as follows:
[0147]
[0148] Step 5: Optimize the reinforcement learning model using the TD3 algorithm, and obtain the solution to the optimization problem based on the optimized reinforcement learning model. This yields the time proportion τ(t) and spatial proportion ξ of RIS deployment for reflecting energy and signal under different time slots t. k (t), such as Figure 2 As shown, specifically:
[0149] (1) Initialize the Actor policy network Critic Review Network And Critic Review Network 2
[0150] and the corresponding Actor target network and Critic target network and Critic target network Where s represents the state, a represents the action, and w represents the network parameter vector;
[0151] (2) The number of training rounds (episode) is initialized to 0;
[0152] (3) The time slot t in the training episode is initialized to 0;
[0153] (4) The Actor policy network outputs action A(t) and obtains a reward value based on the input state S(t). Then, the system transitions to the next state S(t+1) to obtain the training dataset. And store it in an experience replay repository of size B. middle;
[0154] (5) Randomly sample B from the experience replay repository. m A dataset consisting of samples is sent to the Actor policy network, Critic review network 1, Critic review network 2, Actor target network, Critic target network 1, and Critic target network 2; where the i-th sample is denoted as
[0155] (6) The objective function is calculated from the minimum Q-value function of Critic objective network 1 and Critic objective network 2:
[0156]
[0157] Among them, y i γ is the objective function value of the i-th sample; γ is the penalty factor, representing the importance of the expected Q-value function reward in the future. For the bounded smoothed target action, by giving action A i+1 Add bounded noise To obtain:
[0158]
[0159] Among them, A i+1 The output of the Actor target network, To represent a bounded function, Restricted to the lower bound αL and upper bound α H Between, and then output to That is: if but if but if but Formalistically, it can be written as:
[0160]
[0161] Among them, bounded noise Defined as c and -c are bounded functions clip pairs, respectively. The upper and lower bounds; This indicates that the expression follows a pattern with a mean of 0 and a variance of . The normal distribution;
[0162] (7) Critic review network 1 and Critic review network 2 are obtained by minimizing the mean square error between the above objective function and the expected Q value:
[0163]
[0164] Where j = 1, 2 correspond to Critic review network 1 and Critic review network 2, respectively;
[0165] (8) When time interval t passes through T u Next, update the Actor policy network; Actor policy network parameters. The update is obtained by maximizing the expected Q value, and the gradient of the expected Q value is expressed as:
[0166]
[0167] (9) When time interval t passes through T u Next, update the target network; based on the Actor policy network parameters... Critic Review Network 1 Parameter And Critic review network 2 parameters Smoothly update the Actor target network parameters respectively Critic target network 1 parameter And Critic target network 2 parameters
[0168]
[0169] Where ψ∈[0,1] are the smooth update hyperparameters;
[0170] (10) Determine whether the time slot t < T, where T is the total number of time slots; if yes, then t = t + 1 and return to step (4); otherwise, complete the RIS policy generation for one round and continue training for the next round.
[0171] (11) Determine whether the number of rounds episode < EPISODE, where EPISODE is the total number of rounds; if yes, then episode = episode + 1, and return to step (3); otherwise, the optimization ends and the optimized reinforcement learning model is obtained.
[0172] Based on the above example, perform data simulation:
[0173] Assume that in a three-dimensional Cartesian coordinate system, the access point is located at (0,0,10)m and the RIS is located at (5,5,5)m. Within the total number of training time slots T=100, the distance between the industrial robot and the RIS changes uniformly from 5m to 15m with a step size of 0.1.
[0174] Assuming the communication system has an energy transmission efficiency η = 0.7 and a path loss exponent... α = -1.6; Energy and signal balance coefficient δ = 0.5; Path loss κ = -30dB at reference distance d′ = 1m; Additional loss factor for NLoS Maximum transmit power P of the access point max =80W; Minimum signal quality threshold SINR min =12dB; the scale factor of the Poisson distribution is c=0.1.
[0175] In the TD3-based deep reinforcement learning, the Actor policy network, Critic commenting network 1, and Critic commenting network 2 are all composed of fully connected neural networks with three hidden layers, and the optimizer used is the AdamPropOptimizer. The simulation network environment parameters are: a total number of training epochs EPISODE = 500, a total number of training time slots per epoch T = 100, and a random sampling size of B. m =128, the interval T for updating the Actor policy network and target network. u =2, and the learning rates of Actor policy network, Critic comment network 1 and Critic comment network 2 are all set to 0.00001.
[0176] Figure 3 This displays the normalized cumulative reward value corresponding to the total number of training slots in each training round when the total number of training rounds is EPISODE=500. It can be seen that the reward converges as the training time step increases. The TD3 algorithm, based on time and space optimization, achieves a better reward than the three baseline schemes because the TD3 algorithm uses two Critic review networks with an interval T. u =2 updates the Actor policy network, Critic comment network and other strategies, with a more stable learning process. It can learn from the environment and adjust the optimization variables to approach the optimal solution. It also shows that the deployment of RIS plays an important role in improving the wireless power communication of RIS-assisted mobile industrial robots.
[0177] Accordingly, this application also provides an electronic device, comprising: one or more processors; a memory for storing one or more programs; and when the one or more programs are executed by the one or more processors, causing the one or more processors to implement the RIS-assisted mobile industrial robot wireless power-carrying communication optimization method based on the TD3 algorithm described above. Figure 4 The diagram shown is a hardware structure diagram of any device with data processing capabilities for the RIS-assisted mobile industrial robot wireless power-carrying communication optimization method based on the TD3 algorithm provided in this embodiment of the invention, except... Figure 4 In addition to the processor, memory, and network interface shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.
[0178] Accordingly, this application also provides a computer-readable storage medium storing computer instructions, which, when executed by a processor, implement the RIS-assisted mobile industrial robot wireless power-carrying communication optimization method based on the TD3 algorithm described above. The computer-readable storage medium can be an internal storage unit of any data-processing device as described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units of any data-processing device and external storage devices. The computer-readable storage medium is used to store the computer program and other programs and data required by the data-processing device, and can also be used to temporarily store data that has been output or will be output.
[0179] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the disclosure herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and embodiments are to be considered exemplary only.
[0180] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope.
[0181] The above description is merely a preferred embodiment of the present invention. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make many possible variations and modifications to the technical solutions of the present invention using the methods and techniques disclosed above, or modify them into equivalent embodiments with equivalent changes, without departing from the scope of the technical solutions of the present invention. Therefore, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention shall still fall within the protection scope of the technical solutions of the present invention.
Claims
1. A method for optimizing wireless power-carrying communication for RIS-assisted mobile industrial robots based on the TD3 algorithm, characterized in that, The optimization method is based on a RIS-assisted wireless energy-carrying communication system for mobile industrial robots. The communication system includes an access point, a RIS, and industrial robots. The access point has only one antenna with D arrays for beamforming, capable of simultaneously transmitting energy and signal electromagnetic waves. There are K industrial robots, denoted as K. Each is equipped with a single array of antennas; there is only one RIS, which receives energy and electromagnetic waves from the access point and reflects them to K industrial robots; the entire time period is divided into T equal-length time slots, denoted as... The RIS-assisted mobile industrial robot wireless power-carrying communication optimization method based on the TD3 algorithm includes: (1) Model the channels between the access point and the RIS, and the channels between the RIS and the K robots; (2) Model the channel between the access point and the RIS, and the channel path loss between the RIS and the K robots; (3) Based on the channel path loss, establish optimization problems for energy transmission and information transmission under different time slots t, including: The system objective is to optimize the spatial ratio ξ of reflected energy and signal by using the RIS unit. k (t), and the time ratio τ(t) / (1-τ(t)) of the reflected energy alone and the reflected energy to the signal simultaneously, at the minimum threshold SINR for meeting communication quality. min To maximize transmission energy and communication quality, the optimization problem can be formulated as follows: in, Let E be the reward function for the k-th industrial robot in time slot t, and δ be the energy-signal balance coefficient, used to characterize the importance of communication quality in the overall optimization objective; k (t) represents the energy received by the k-th industrial robot in time slot t; SINR k (t) represents the signal-to-interference-plus-noise ratio (SINR) of the k-th industrial robot in time slot t; C1 is the minimum threshold for communication quality, i.e., the minimum SINR constraint. min C2 represents the finite-time constraint; C3 represents the maximum energy transfer power constraint p. max V k EH C4 is the precoded vector used by the access point to transmit energy; C4 is the RIS ratio constraint. (4) Establish a deep reinforcement learning model based on the channel environment, the variables to be optimized, and the target to be optimized; (5) The reinforcement learning model is optimized using the TD3 algorithm, and the solution to the optimization problem is obtained based on the optimized reinforcement learning model. The time and space ratios of RIS deployment for reflecting energy and signal are obtained under different time slots t.
2. The RIS-assisted mobile industrial robot wireless power-carrying communication optimization method based on the TD3 algorithm according to claim 1, characterized in that, Modeling the channels between the access point and the RIS, and between the RIS and the K robots, includes: (1) Establish a three-dimensional Cartesian coordinate system for all communication nodes; (2) RIS is equipped with m rows and n columns, total There are 1 reflective units, and the array of reflective units is denoted as . The reflection unit in the i-th row and j-th column is denoted as Where i∈{1,2,…,m},j∈{1,2,…,n}, the amplitude and phase of each reflection unit can be adjusted; (3) Divide the entire time period into T equal-length time slots, denoted as . Each time slot t is divided into two stages: an energy transmission stage and an information transmission stage. In the energy transmission stage, all RIS units are used for energy transmission, and their time allocation within time slot t is τ(t), where 0 ≤ τ(t) ≤ 1. In the information transmission stage, for the industrial robot k, only… Each RIS unit is used for transmitting information, denoted as . The other part of the RIS units are used for energy transmission; the time of the information transmission phase accounts for 1-τ(t) in time slot t; the ratio of the number of RIS units used for information transmission to the total number of RIS units in the information transmission phase is denoted as τ(t). (4) The beamforming at the access point employs a linear precoding strategy, and its transmitted signal is represented as follows: in, and These are electromagnetic waves for energy transmission and electromagnetic waves for information transmission, respectively. and These are the precoded vectors used by the access point to transmit energy and information, respectively. Represents a complex vector in a D×1 space; and The original signals used by the access point for transmitting energy and information are respectively, both following a cyclic symmetric complex Gaussian distribution with mean 0 and variance 1; the total power p used by the access point for transmitting energy is expressed as: in, represents the conjugate transpose, and ||·|| represents the 2-norm of the vector; (5) In each time slot t, the industrial robot Energy received E k (t) is represented as: Among them, Z k From the access point to the RIS unit The baseband equivalent channel; From RIS unit The baseband equivalent channel to industrial robot k; It is a RIS unit The diagonal matrix of the reflection coefficient matrix, where θ l and β l These represent the phase and amplitude of the RIS unit l, respectively; From the access point to part of the RIS unit The baseband equivalent channel; From a portion of the RIS unit The baseband equivalent channel to industrial robot k; It is a part of the RIS unit The diagonal matrix of the reflection coefficient matrix; η∈(0,1) is the energy transfer efficiency; ω i,j;k (t) is an indicator function that indicates the RIS unit. Whether it is used for transmitting information is indicated as follows: (6) During the energy transfer phase in each time slot t, where the proportion is τ(t) and 0≤τ(t)≤1, all RIS units They are evenly distributed to each industrial robot k; (7) In each time slot t, the industrial robot The received information y k (t) is represented as: Among them, v k The noise level of each industrial robot is represented by the noise power formula. Additive white Gaussian noise; (8) Considering both line-of-sight and non-line-of-sight channels, the channel model from RIS to industrial robot is modeled as following a Rayleigh distribution.
3. The RIS-assisted mobile industrial robot wireless power-carrying communication optimization method based on the TD3 algorithm according to claim 1, characterized in that, The modeling of the channel between the access point and the RIS, and the channel path loss between the RIS and the K robots includes: (1) The channel between the access point and the RIS is mainly a line-of-sight (LoS) channel. Channel path loss between The model is as follows: Among them, (x i,j ,y i,j ,z i,j ) represents the RIS unit in the i-th row and j-th column. The location in the three-dimensional Cartesian coordinate system; the access point is at the origin of the three-dimensional Cartesian coordinate system; α represents the distance from the access point to the RIS unit. The channel path loss exponent; κ represents the channel path loss at the reference distance d′; (2) The access point is assumed to be a point mass in three-dimensional space; (3) The channel between the RIS and the industrial robot includes a line-of-sight channel and a non-line-of-sight (NLoS) channel. Channel path loss L RI The model is as follows: in, These represent the probabilities of the LoS channel and the NLoS channel between the RIS and the industrial robot, respectively. This represents the channel path loss index from the RIS to the industrial robot; Indicates the additional path loss factor under NLoS channel; d k (t) represents the distance from RIS to each industrial robot k in each time slot t; (4) In the channel between the RIS and the industrial robot, the d in each time slot t k (t) is considered a constant; (5) The LoS channel between the RIS and the industrial robot is modeled as follows: Wherein, variable X represents the number of times the channel between the RIS and the target industrial robot is blocked by other mobile industrial robots per unit time, and X follows a Poisson distribution with parameter λ; the probability distribution function of X is expressed as: Where λ is proportional to unit time, denoted as: Where c, c>0 is a scaling factor used to represent the severity of the blockage; (6) The NLoS channel between the RIS and the industrial robot is modeled as follows: (7) Considering radio frequency interference, the signal-to-interference-plus-noise ratio (SINR) is used to represent the industrial robot. The quality of the received signal is defined as: Among them, SINR k (t) represents the signal-to-interference-plus-noise ratio (SIR) of the k-th industrial robot in time slot t.
4. The RIS-assisted mobile industrial robot wireless power-carrying communication optimization method based on the TD3 algorithm according to claim 1, characterized in that, The process of establishing a deep reinforcement learning model based on the channel environment, the variables to be optimized, and the objective to be optimized includes: (1) The state space of the deep reinforcement learning model includes: the transmitted signal G of the access point; the distance d between RIS and the k-th industrial robot. k (t); The baseband equivalent channel h of RIS and the k-th industrial robot k (t); (2) The action space output by the deep reinforcement learning model includes: the time proportion τ(t) and the spatial proportion ξ at time slot t. k (t); (3) The reward function of the deep reinforcement learning model is modeled as: Optimization objective (4) The optimization objective The received dryness ratio must not be lower than the minimum threshold SINR. min Due to the constraints, the reward value of the k-th industrial robot in time slot t is... The model is as follows:
5. The RIS-assisted mobile industrial robot wireless power-carrying communication optimization method based on the TD3 algorithm according to claim 1, characterized in that, In step (5), the reinforcement learning model is optimized using the TD3 algorithm, and the solution to the optimization problem is obtained based on the optimized reinforcement learning model. The time ratio τ(t) and spatial ratio ξ of RIS deployment for reflecting energy and signal are obtained under different time slots t. k (t), including: (1) Initialize the Actor policy network Critic Review Network And Critic Review Network 2 and the corresponding Actor target network and Critic target network and Critic target network Where s represents the state, a represents the action, and w represents the network parameter vector; (2) The number of training rounds (episode) is initialized to 0; (3) The time slot t in the training episode is initialized to 0; (4) The Actor policy network outputs action A(t) and obtains a reward value based on the input state S(t). Then, the system transitions to the next state S(t+1) to obtain the training dataset. And store it in an experience replay repository of size B. middle; (5) Randomly sample B from the experience replay repository. m A dataset consisting of samples is sent to the Actor policy network, Critic review network 1, Critic review network 2, Actor target network, Critic target network 1, and Critic target network 2; where the i-th sample is denoted as (6) The objective function is calculated from the minimum Q-value function of Critic objective network 1 and Critic objective network 2: Among them, y i γ is the objective function value of the i-th sample; γ is the penalty factor, representing the importance of the expected Q-value function reward in the future. For the bounded smoothed target action, by giving action A i+1 Add bounded noise To obtain; (7) Critic review network 1 and Critic review network 2 are obtained by minimizing the mean square error between the above objective function and the expected Q value: Where j = 1, 2 correspond to Critic review network 1 and Critic review network 2, respectively; (8) When time interval t passes through T u Next, update the Actor policy network; Actor policy network parameters. The update is obtained by maximizing the expected Q value, and the gradient of the expected Q value is expressed as: (9) When time interval t passes through T u Next, update the target network; based on the Actor policy network parameters... Critic Review Network 1 Parameter And Critic review network 2 parameters Smoothly update the Actor target network parameters respectively Critic target network 1 parameter And Critic target network 2 parameters Where ψ∈[0,1] are the smooth update hyperparameters; (10) Determine whether the time slot t < T is satisfied, where T is the total number of time slots; if so, then t = t + 1, and return to step (4); Otherwise, after completing one round of RIS policy generation, continue with the next round of training; (11) Determine whether the number of rounds episode < EPISODE, where EPISODE is the total number of rounds; if yes, then episode = episode + 1, and return to step (3); otherwise, the optimization ends and the optimized reinforcement learning model is obtained.
6. The RIS-assisted mobile industrial robot wireless power-carrying communication optimization method based on the TD3 algorithm according to claim 5, characterized in that, Bounded smoothed target action By giving action A i+1 Add bounded noise To obtain: Among them, A i+1 The output of the Actor target network, To represent a bounded function, Restricted to the lower bound α L and upper bound α H Between, and then output to That is: if but if but if but Represented as: Among them, bounded noise Defined as c and -c are bounded functions clip pairs, respectively. The upper and lower bounds; This indicates that the expression follows a pattern with a mean of 0 and a variance of . It follows a normal distribution.
7. An electronic device comprising a memory and a processor, characterized in that, The memory is coupled to the processor; wherein the memory is used to store program data, and the processor is used to execute the program data to implement the RIS-assisted mobile industrial robot wireless power-carrying communication optimization method based on the TD3 algorithm as described in any one of claims 1-6.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the RIS-assisted mobile industrial robot wireless power-carrying communication optimization method based on the TD3 algorithm as described in any one of claims 1-6.
9. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instruction is executed by the processor, it implements the RIS-assisted mobile industrial robot wireless power-carrying communication optimization method based on the TD3 algorithm as described in any one of claims 1-6.
Citation Information
Patent Citations
DDPG-based IRS-assisted cognitive radio system beam forming method
CN117767987A
Unmanned aerial vehicle secrecy rate optimization method based on evolutionary reinforcement learning
CN118921654A