TD3 algorithm-based RIS-assisted mobile industrial robot simultaneous wireless information and power transfer optimization method

Through the TD3 algorithm, the RIS reflected energy and signal ratio are optimized, signal blocking and battery capacity limitations in wireless industrial robot communication are solved, and communication quality and battery life are improved.

CN120302316AActive Publication Date: 2025-07-11ZHEJIANG UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510381181.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-11
Estimated Expiration
2045-03-28

AI Technical Summary

Technical Problem

无线工业机器人通信面临信号阻塞和电池容量限制的问题,导致通信质量差和续航时间短。

Method used

The RIS-assisted mobile industrial robot wireless portable communication method based on the TD3 algorithm is adopted to optimize the spatial and time ratio of RIS reflected energy and signal by modeling channel and path loss, and maximize the energy and communication quality.

Benefits of technology

While meeting the minimum communication quality threshold, the communication quality and battery life of wireless industrial robots are optimized, solving the problems of signal blocking and battery capacity limitation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120302316A_ABST
    Figure CN120302316A_ABST
Patent Text Reader

Abstract

The invention discloses a TD3 algorithm-based RIS-assisted mobile industrial robot simultaneous wireless information and power transfer optimization method. According to the method, on the basis of the position change and the dynamic shielding condition of the mobile industrial robot, under the condition that the minimum threshold of communication quality is met, an intelligent reflection surface (RIS) is deployed for reflecting energy and the space proportion of signals, pure reflection energy and the time proportion of the reflection energy and the signals, and by means of an intelligent decision of TD3 (Twin Delayed Development Policy Gradient), the communication quality of the mobile industrial robot is improved, so that the communication quality of the mobile industrial robot is improved, and the communication quality of the mobile industrial robot is improved. And maximization of transmission energy and communication quality is realized. The problems that the industrial robot is poor in wireless communication quality and short in endurance time are solved at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of communications, and specifically relates to an optimization method for RIS-assisted wireless energy-harvesting communication of mobile industrial robots based on the TD3 algorithm. Background Art

[0002] Wireless communication technology and robotics have brought many conveniences to Industry 4.0. Compared with wired communication, industrial wireless communication is easy to deploy and maintain, and can reduce costs by 90%; and it can achieve free movement of on-site equipment, flexible network structures, and diverse applications. At the same time, wireless communication-assisted mobile industrial robots can be clustered without being restricted by the working space, realizing diverse manufacturing and adapting to different environments, which helps to reduce costs.

[0003] However, the application of wireless industrial robots faces two major problems. First, critical tasks in industrial production require deterministic communication. Due to moving objects (such as industrial robots, trucks, robotic arms, etc.) and densely deployed metal machinery, wireless signals are more vulnerable to blockage. In addition, battery capacity also limits the performance of industrial robots, such as working range, service time, payload, etc.

[0004] Existing work has not solved the above two major problems. Therefore, how to deploy RIS for the spatial ratio of reflected energy and signals, as well as the time ratio of simply reflecting energy and simultaneously reflecting energy and signals, according to the position changes and dynamic occlusion conditions of mobile industrial robots, and with the intelligent decision-making of the TD3 reinforcement learning model, to maximize the transmitted energy and communication quality while meeting the minimum threshold of communication quality has become a very important issue. Summary of the Invention

[0005] The purpose of the present invention is to provide an optimization method for RIS-assisted wireless energy-harvesting communication of mobile industrial robots based on the TD3 algorithm in view of the deficiencies of the prior art.

[0006] The purpose of the present invention is achieved through the following technical solutions: An optimization method for RIS-assisted wireless energy-harvesting communication of mobile industrial robots based on the TD3 algorithm, the optimization method being based on a RIS-assisted wireless energy-harvesting communication system for mobile industrial robots; the communication system includes an access point, a RIS, and industrial robots; there is only one antenna with D arrays for beamforming on the access point, which can simultaneously transmit energy and signal electromagnetic waves; there are K industrial robots, denoted as Each is equipped with an antenna with a single array; there is only one RIS, which receives the energy and signal electromagnetic waves from the access point and reflects them to the K industrial robots; the RIS-assisted wireless energy-harvesting communication optimization method based on the TD3 algorithm includes:

[0007] (1) Model the channels between the access point and the RIS, and between the RIS and the K robots.

[0008] (2) Model the path losses of the channels between the access point and the RIS, and between the RIS and the K robots.

[0009] (3) Based on the path losses of the channels, establish optimization problems regarding energy transfer and information transfer at different time slots t.

[0010] (4) Establish a deep reinforcement learning model according to the channel environment, variables to be optimized, and objectives to be optimized.

[0011] (5) Use the TD3 algorithm to optimize the reinforcement learning model, and obtain the solutions to the optimization problems according to the optimized reinforcement learning model, so as to obtain the time ratios and spatial ratios for deploying the RIS to reflect energy and signals at different time slots t.

[0012] Furthermore, modeling the channels between the access point and the RIS, and between the RIS and the K robots includes:

[0013] (1) Establish a three-dimensional Cartesian coordinate system for all communication nodes.

[0014] (2) The RIS is equipped with m rows and n columns, a total of reflecting units, and the reflecting unit array is denoted as The reflecting unit in the i-th row and j-th column is denoted as where i ∈ {1, 2, …, m}, j ∈ {1, 2, …, n}, and the amplitude and phase of each reflecting unit can be adjusted.

[0015] (3) Divide the entire time period into T equal-length time slots, denoted as where each time slot t is divided into two stages, an energy transfer stage and an information transfer stage; in the energy transfer stage, all RIS units are used to transfer energy, and its time accounts for τ(t) in the time slot t, 0 ≤ τ(t) ≤ 1; in the information transfer stage, for industrial robot k, only RIS units are used to transfer information, denoted as while the other part of the RIS units is used to transfer energy; the time of the information transfer stage accounts for 1 - τ(t) in the time slot t; denote the ratio of the number of RIS units used to transfer information in the information transfer stage to the total number of RIS units as

[0016] (4) The beamforming of the access point adopts a linear precoding strategy, and its transmitted signal is expressed as:

[0017]

[0018] Among them, and are the electromagnetic waves for energy transmission and information transmission respectively; and are the precoding vectors used by the access point for energy transmission and information transmission respectively; represents a complex vector in the D×1 space; and are the original signals used by the access point for energy transmission and information transmission respectively, both of which follow a circularly symmetric complex Gaussian distribution with a mean of 0 and a variance of 1; the total power p used by the access point for energy transmission is expressed as:

[0019]

[0020] Among them, represents conjugate transpose, and ||·|| represents the 2-norm of a vector;

[0021] (5) In each time slot t, the energy E received by the industrial robot k (t) is expressed as:

[0022]

[0023] Among them, Z k is the baseband equivalent channel from the access point to the RIS unit ; is the baseband equivalent channel from the RIS unit to the industrial robot k; is the diagonal matrix of the reflection coefficient matrix of the RIS unit , where θ l and β l represent the phase and amplitude of the RIS unit l respectively; is the baseband equivalent channel from the access point to some RIS units ; is the baseband equivalent channel from some RIS units to the industrial robot k; is the diagonal matrix of the reflection coefficient matrix of some RIS units ; η∈(0,1) is the energy transmission efficiency; ω i,j;k (t) is an indicator function indicating whether the RIS unit is used for information transmission, and is expressed as:

[0024]

[0025] (6) In the energy transmission phase with a proportion of τ(t), 0≤τ(t)≤1 in each time slot t, all RIS units is evenly allocated to each industrial robot k;

[0026] (7) In each time slot t, the industrial robot the received information y k (t) is expressed as:

[0027]

[0028] where v k represents the noise of each industrial robot, which follows additive white Gaussian noise with a noise power of ;

[0029] (8) Considering both the line-of-sight and non-line-of-sight channels simultaneously, the channel from the RIS to the industrial robot is modeled as following a Rayleigh distribution.

[0030] Furthermore, the path loss modeling of the channel between the access point and the RIS, and the channels between the RIS and the K robots includes:

[0031] (1) The channel between the access point and the RIS is mainly a line-of-sight (LoS) channel, and the path loss of the channel between the access point and the RIS unit is modeled as:

[0032]

[0033] where (x i,j , y i,j , z i,j ) represents the position of the RIS unit in the i-th row and j-th column in the three-dimensional Cartesian coordinate system; the access point is at the origin position of the three-dimensional Cartesian coordinate system; α represents the path loss exponent of the channel from the access point to the RIS unit ; κ represents the path loss at the reference distance d'; ;

[0034] (2) The access point is assumed to be a point mass in three-dimensional space;

[0035] (3) The channel between the RIS and the industrial robot includes a line-of-sight channel and a non-line-of-sight (NLoS) channel. The path loss L between the RIS and the industrial robot RI is modeled as:

[0036]

[0037] where respectively represent the probabilities of the LoS channel and the NLoS channel between the RIS and the industrial robot; Denotes the path loss exponent of the communication path from the RIS to the industrial robot; Denotes the additional path loss factor under the NLoS channel; d k (t) is the distance from the RIS to each industrial robot k at each time slot t;

[0038] In the channel between the RIS and the industrial robot described in (4), d within each time slot t k (t) is considered a constant;

[0039] (5) Model the LoS channel between the RIS and the industrial robot as:

[0040]

[0041] where the variable X represents the number of times the channel between the RIS and the target industrial robot is blocked by other mobile industrial robots per unit time, and X follows a Poisson distribution with parameter λ; the probability distribution function of X is expressed as:

[0042]

[0043] where λ is proportional to the unit time, denoted as:

[0044]

[0045] where c, c > 0 is a proportionality factor used to represent the severity of the blockage;

[0046] (6) Model the NLoS channel between the RIS and the industrial robot as:

[0047]

[0048] (7) Considering radio frequency interference, use the signal-to-interference-plus-noise ratio to represent the signal quality received by the industrial robot is defined as:

[0049]

[0050] where SINR k (t) is the signal-to-interference-plus-noise ratio of the k-th industrial robot at time slot t.

[0051] Furthermore, the optimization problems regarding energy transmission and signal transmission at different time slots t established according to the communication path loss include:

[0052] The system objective is to optimize the spatial ratio ξ of the RIS unit for reflecting energy and signals k(t), as well as the simple reflection energy and the time ratio τ(t) / (1 - τ(t)) of the reflected energy and the signal, to maximize the transmission energy and communication quality under the minimum threshold SINR of the communication quality min under which, the optimization problem is formulated as:

[0053]

[0054] wherein, is the reward function of the kth industrial robot at time slot t, δ is the energy and signal balance coefficient, which is used to characterize the importance of communication quality in the whole optimization objective; C1 is the minimum threshold of communication quality, that is, the minimum signal-to-interference-plus-noise ratio constraint SINR min ; C2 is the finite time constraint; C3 is the maximum energy transmission power constraint p max ; C4 is the RIS ratio constraint.

[0055] Furthermore, the establishment of the deep reinforcement learning model according to the channel environment, variables to be optimized, and objectives to be optimized includes:

[0056] (1) The state space of the deep reinforcement learning model includes: the transmitted signal G of the access point; the distance d k (t) between the RIS and the kth industrial robot; the baseband equivalent channel h k (t) between the RIS and the kth industrial robot;

[0057] (2) The action space of the deep reinforcement learning model includes: the time ratio and the spatial ratio ξ k (t) at time slot t;

[0058] (3) Modeling the reward function of the deep reinforcement learning model as: the optimization objective described in claim 4

[0059] (4) The optimization objective The received dry ratio must not be lower than the minimum threshold SINR min constraint, so the reward value of the kth industrial robot at time slot t is modeled as:

[0060]

[0061] Furthermore, in step (5), the TD3 algorithm is used to optimize the reinforcement learning model, and the solution of the optimization problem is obtained according to the optimized reinforcement learning model, and the time ratio τ(t) and the spatial ratio ξ k (t) for deploying the RIS to reflect energy and signal at different time slots t are obtained, including:

[0062] (1) Initializing the Actor policy network Critic Critic network and Critic Critic network 2

[0063] and the corresponding Actor target network and Critic target network and Critic target network where s represents the state, a represents the action, and w represents the network parameter vector;

[0064] (2) The number of training episodes episode is initialized to 0;

[0065] (3) The time slot t in the training episode episode is initialized to 0;

[0066] (4) The Actor policy network outputs the action A(t) according to the input state S(t) and obtains the reward value and then transitions to the next state S(t + 1), obtaining the training data set and stores it in the experience replay repository of size B ;

[0067] (5) Randomly sample B m samples from the experience replay repository to form a data set and send it to the Actor policy network, Critic Critic network 1, Critic Critic network 2, Actor target network, Critic target network 1, and Critic target network 2; among them, the i-th sample is denoted as

[0068] (6) The objective function is calculated by the minimum Q-value function of Critic target network 1 and Critic target network 2:

[0069]

[0070] where y i is the objective function value of the i-th sample; γ is the penalty factor, indicating the importance of the future expected Q-value function reward; is the bounded smoothed target action, obtained by adding bounded noise i+1 to the action A ;

[0071] (7) Critic Critic network 1 and Critic Critic network 2 are obtained by minimizing the mean square error between the above objective function and the expected Q-value:

[0072]

[0073] where \(j = 1, 2\) corresponds to Critic comment network 1 and Critic comment network 2 respectively;

[0074] (8) When the time slot \(t\) passes \(T\) u times, update the Actor policy network; the update of the Actor policy network parameters is obtained by maximizing the expected Q value, and the gradient of the above expected Q value is expressed as:

[0075]

[0076] (9) When the time slot \(t\) passes \(T\) u times, update the target network; according to the Actor policy network parameters Critic comment network 1 parameters and Critic comment network 2 parameters smoothly update the Actor target network parameters Critic target network 1 parameters and Critic target network 2 parameters

[0077]

[0078] where \(\psi\in[0, 1]\) is the smooth update hyperparameter;

[0079] (10) Determine whether \(t < T\) is satisfied, where \(T\) is the total number of time slots; if so, \(t=t + 1\), and return to step (4); otherwise, complete the generation of the RIS policy for one round of episodes, and continue the training for the next round of episodes;

[0080] (11) Determine whether the number of episodes \(episode < EPISODE\) is satisfied, where \(EPISODE\) is the total number of episodes; if so, \(episode=episode + 1\), and return to step (3); otherwise, the optimization ends, and the optimized reinforcement learning model is obtained.

[0081] Furthermore, the bounded and smoothed target action is obtained by adding bounded noise i+1 to the action \(A\) as follows:

[0082]

[0083] where \(A\) i+1 is the output of the Actor target network, represents a bounded function that restricts to the lower bound \(\alpha\) L and the upper bound \(\alpha\) Hbetween, and then output to That is: If then If then If then It is expressed as:

[0084]

[0085] Among them, bounded noise is defined as c and -c are respectively the upper and lower bounds of the bounded function clip for ; denotes a normal distribution with a mean of 0 and a variance of ;

[0086] The present invention also provides an electronic device, including a memory and a processor, the memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned RIS-assisted mobile industrial robot wireless energy-harvesting communication optimization method based on the TD3 algorithm.

[0087] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the above-mentioned RIS-assisted mobile industrial robot wireless energy-harvesting communication optimization method based on the TD3 algorithm.

[0088] The present invention also provides a computer program product, including computer program / instructions, characterized in that when the computer program / instructions are executed by a processor, they implement the above-mentioned RIS-assisted mobile industrial robot wireless energy-harvesting communication optimization method based on the TD3 algorithm.

[0089] The beneficial effects of the present invention are as follows: The present invention provides a RIS-assisted mobile industrial robot wireless energy-harvesting communication optimization method based on the TD3 algorithm. According to the position change and dynamic occlusion of the mobile industrial robot, under the condition of meeting the minimum threshold of communication quality, by deploying the spatial ratio of RIS for reflecting energy and signals, and the time ratio of simply reflecting energy and reflecting energy and signals simultaneously, with the help of the intelligent decision-making of the TD3 reinforcement learning model, the maximization of transmission energy and communication quality is achieved. The present invention also solves the problems of poor wireless communication quality and short battery life of industrial robots. Description of the Drawings

[0090] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for use in the embodiments or the description of the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0091] Figure 1 It is a diagram of the RIS-assisted wireless energy-harvesting communication model of the mobile industrial robot provided by the embodiment of the present invention;

[0092] Figure 2 It is a framework diagram of the TD3 algorithm provided by the embodiment of the present invention;

[0093] Figure 3 It is a reward diagram of the TD3 algorithm under the number of training steps provided by the embodiment of the present invention;

[0094] Figure 4 It is a schematic structural diagram of the electronic device provided by the embodiment of the present invention. Detailed implementation manners

[0095] To better understand the technical solutions of the present application, the following will describe the embodiments of the present application in detail with reference to the accompanying drawings.

[0096] It should be clear that the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.

[0097] The RIS-assisted mobile industrial robot wireless energy-harvesting communication optimization method provided by the present invention, according to the position change and dynamic occlusion situation of the mobile industrial robot, under the condition of meeting the minimum threshold of communication quality, realizes the maximization of transmitted energy and communication quality through deploying the spatial ratio of RIS for reflecting energy and signals, as well as the time ratio of simply reflecting energy and simultaneously reflecting energy and signals, with the help of the intelligent decision-making of the TD3 reinforcement learning model. The present invention also solves the problems of poor wireless communication quality and short battery life of industrial robots.

[0098] This optimization method is based on the RIS-assisted mobile industrial robot wireless energy-harvesting communication system as shown in Figure 1 ; the communication system includes an access point, RIS, and industrial robots; there is only one antenna with D arrays for beamforming on the access point, which can simultaneously transmit energy and signal electromagnetic waves; there are K industrial robots, denoted as Antennas with a single array are installed respectively; there is only one RIS, which receives energy and signal electromagnetic waves from the access point and reflects them to K industrial robots; the RIS-assisted wireless energy-harvesting communication optimization method for mobile industrial robots based on the TD3 algorithm includes:

[0099] Step 1: Model the channels between the access point and the RIS, and between the RIS and the K robots; specifically:

[0100] (1) Establish a three-dimensional Cartesian coordinate system for all communication nodes;

[0101] (2) The RIS is equipped with m rows and n columns, a total of reflective units, and the reflective unit array is denoted as The reflective unit in the i-th row and j-th column is denoted as where i ∈ {1, 2, …, m}, j ∈ {1, 2, …, n}, and the amplitude and phase of each reflective unit can be adjusted;

[0102] (3) Divide the entire time period into T equal-length time slots, denoted as where each time slot t is divided into two stages, the energy transmission stage and the information transmission stage; in the energy transmission stage, all RIS units are used to transmit energy, and its time accounts for τ(t) in time slot t, 0 ≤ τ(t) ≤ 1; in the information transmission stage, for industrial robot k, only RIS units are used to transmit information, denoted as while the other part of the RIS units is used to transmit energy; the time of the information transmission stage accounts for 1 - τ(t) in time slot t; in addition, the ratio of the number of RIS units used to transmit information in the information transmission stage to the total number of RIS units is denoted as

[0103] (4) The beamforming of the access point adopts a linear precoding strategy, and its transmitted signal is expressed as:

[0104]

[0105] where, and are the energy transmission electromagnetic wave and the information transmission electromagnetic wave respectively; and are the precoding vectors used by the access point to transmit energy and transmit information respectively; represents a complex vector in the D×1 space; and are the original signals used by the access point to transmit energy and transmit information respectively, both of which follow a circularly symmetric complex Gaussian distribution with a mean of 0 and a variance of 1, that is: Therefore, there is a mean Furthermore, the total power p used by the access point to transmit energy is expressed as:

[0106]

[0107] where denotes the conjugate transpose, and ||·|| represents the 2-norm of a vector;

[0108] (5) In each time slot t, the energy E received by the industrial robot k (t) is expressed as:

[0109]

[0110] where Z k is the baseband equivalent channel from the access point to the RIS unit ; is the baseband equivalent channel from the RIS unit to the industrial robot k; is the diagonal matrix of the reflection coefficient matrix of the RIS unit , where θ l and β l represent the phase and amplitude of the RIS unit l, respectively; similarly, is the baseband equivalent channel from the access point to some of the RIS units ; is the baseband equivalent channel from some of the RIS units to the industrial robot k; is the diagonal matrix of the reflection coefficient matrix of some of the RIS units ; η ∈ (0, 1) is the energy transfer efficiency; ω i,j;k (t) = 0 represents that the RIS unit is used to transmit information, and vice versa. Therefore, the indicator function ω i,j;k (t) is expressed as:

[0111]

[0112] (6) In the energy transfer phase with a proportion of τ(t), 0 ≤ τ(t) ≤ 1, in each time slot t, all the RIS units are evenly allocated to each industrial robot k;

[0113] (7) In each time slot t, the information y received by the industrial robot k (t) is expressed as:

[0114]

[0115] where v kRepresents the noise of each industrial robot, and the noise power is The additive Gaussian white noise is denoted by

[0116] (8) Considering both line-of-sight and non-line-of-sight channels, the channel from RIS to the industrial robot is modeled as obeying the Rayleigh distribution.

[0117] Step 2: Model the channel path loss between the access point and RIS, and between RIS and K robots; specifically:

[0118] (1) The channel between the access point and the RIS is mainly a line-of-sight (LoS) channel. The channel path loss between Modeled as:

[0119]

[0120] Among them, (x i,j ,y i,j ,z i,j ) represents the RIS unit in row i and column j The position of the access point in the three-dimensional Cartesian coordinate system; the origin of the access point in the three-dimensional Cartesian coordinate system; α represents the distance from the access point to the RIS unit The channel path loss index; κ represents the channel path loss at the reference distance d′=1m;

[0121] (2) The access point is assumed to be a mass point in three-dimensional space;

[0122] (3) The channel between the RIS and the industrial robot includes a line-of-sight (LoS) channel and a non-line-of-sight (NLoS) channel. The channel path loss L RI Modeled as:

[0123]

[0124] in, represent the probabilities of the LoS channel and the NLoS channel between the RIS and the industrial robot respectively; represents the channel path loss index from RIS to the industrial robot; represents the additional path loss factor under NLoS channel; d k (t) is the distance from RIS to each industrial robot k at each time slot t;

[0125] (4) In the channel between the RIS and the industrial robot, d(t) within each time slot t is considered a constant because the time slot t has a very short period during which the change in d(t) is very small and can be ignored; k (5) Model the LoS channel between the RIS and the industrial robot as: k (5) Model the LoS channel between the RIS and the industrial robot as:

[0126] (5) Model the LoS channel between the RIS and the industrial robot as:

[0127]

[0128] where the variable X represents the number of times the channel between the RIS and the target industrial robot is blocked by other moving industrial robots per unit time, and X follows a Poisson distribution with parameter λ; the probability distribution function of X is expressed as:

[0129]

[0130] where λ is proportional to the unit time and is denoted as:

[0131]

[0132] where c, c > 0 is a proportionality factor used to represent the severity of the blockage, which is usually proportional to the moving speed, quantity, size, etc. of the industrial robots;

[0133] (6) Model the NLoS channel between the RIS and the industrial robot as:

[0134]

[0135] (7) Considering radio frequency interference, use the signal-to-interference-plus-noise ratio to represent the signal quality received by the industrial robot and define it as:

[0136]

[0137] where SINR k (t) is the signal-to-interference-plus-noise ratio of the k-th industrial robot at time slot t.

[0138] Step 3: Establish an optimization problem regarding energy transmission and information transmission at different time slots t according to the path loss of the channel;

[0139] Specifically, the system objective is to maximize the transmission energy and communication quality by optimizing the spatial ratio ξ(t) of the RIS unit for reflecting energy and signals, and the time ratio τ(t) / (1 - τ(t)) of simply reflecting energy and reflecting energy and signals simultaneously, while meeting the minimum threshold SINR k of the communication quality. This optimization problem is formulated as: min under which the transmission energy and communication quality are maximized. This optimization problem is formulated as:

[0140]

[0141] wherein, is the reward function of the k-th industrial robot at time slot t, δ is the energy and signal balance coefficient, which is a positive number and is used to characterize the importance of communication quality in the entire optimization objective; C1 is the minimum threshold of communication quality, that is, the minimum signal-to-interference-plus-noise ratio constraint SINR min ; C2 is the finite time constraint; C3 is the maximum energy transmission power constraint p max ; C4 is the RIS ratio constraint. The above optimization problem contains discrete variable ξ k (t) and continuous variable τ(t), so this optimization problem is a mixed-integer non-linear optimization problem. There is a problem of state explosion in all possible solutions of this optimization problem, and the non-convex constraint C4 makes it difficult to solve.

[0142] Step 4: Establish a deep reinforcement learning model according to the channel environment, variables to be optimized, and objectives to be optimized; specifically:

[0143] (1) The state space of the deep reinforcement learning model includes: the transmitted signal G of the access point; the distance d k (t) between the RIS and the k-th industrial robot; the baseband equivalent channel h k (t) between the RIS and the k-th industrial robot;

[0144] (2) The action space of the deep reinforcement learning model includes: the time ratio and space ratio ξ k (t) at time slot t;

[0145] (3) Model the reward function of the deep reinforcement learning model as: the optimization objective described in Step 3

[0146] (4) The optimization objective The received dry ratio must not be lower than the minimum threshold SINR min constraint, so the reward value of the k-th industrial robot at time slot t is modeled as:

[0147]

[0148] Step 5: Optimize the reinforcement learning model using the TD3 algorithm, and obtain the solution of the optimization problem according to the optimized reinforcement learning model, to obtain the time ratio τ(t) and space ratio ξ k (t) for deploying the RIS to reflect energy and signal at different time slots t, as Figure 2 shown, specifically:

[0149] (1) Initialize the Actor policy network Critic Critic network and Critic Critic network 2

[0150] and the corresponding Actor target network and Critic target network and Critic target network where s represents the state, a represents the action, and w represents the network parameter vector;

[0151] (2) The number of training episodes episode is initialized to 0;

[0152] (3) The time slot t in the training episode episode is initialized to 0;

[0153] (4) The Actor policy network outputs the action A(t) according to the input state S(t) and obtains the reward value and then transitions to the next state S(t + 1), obtaining the training data set and stores it in the experience replay repository of size B ;

[0154] (5) Randomly sample B m samples from the experience replay repository to form a data set, and send it to the Actor policy network, Critic Critic network 1, Critic Critic network 2, Actor target network, Critic target network 1, and Critic target network 2; where the i-th sample is denoted as

[0155] (6) The objective function is calculated from the minimum Q-value functions of Critic target network 1 and Critic target network 2:

[0156]

[0157] where y i is the objective function value of the i-th sample; γ is the penalty factor, representing the importance of the future expected Q-value function reward; is the bounded and smoothed target action, obtained by adding bounded noise i+1 to the action A :

[0158]

[0159] where A i+1 is the output of the Actor target network, represents the bounded function, and bounds to the lower bound αL and the upper bound α H and then output to That is: If then If then If then Formally, it can be written as:

[0160]

[0161] where the bounded noise is defined as c and -c are the upper and lower bounds of the bounded function clip for respectively; denotes a normal distribution with mean 0 and variance ;

[0162] (7) The Critic networks 1 and 2 are obtained by minimizing the mean square error between the above objective function and the expected Q value:

[0163]

[0164] where j = 1, 2 correspond to the Critic networks 1 and 2 respectively;

[0165] (8) When the time slot t passes T u times, the Actor policy network is updated; the update of the Actor policy network parameters is obtained by maximizing the expected Q value, and the gradient of the above expected Q value is expressed as:

[0166]

[0167] (9) When the time slot t passes T u times, the target network is updated; according to the Actor policy network parameters the Critic network 1 parameters and the Critic network 2 parameters the Actor target network parameters the Critic target network 1 parameters and the Critic target network 2 parameters

[0168]

[0169] where ψ ∈ [0, 1] is the smoothing update hyperparameter;

[0170] (10) Determine whether it satisfies \(t < T\), where \(T\) is the total number of time slots; if so, then \(t=t + 1\), and return to step (4); otherwise, complete the generation of the RIS strategy for one round of episode number, and continue the training for the next round of episode number;

[0171] (11) Determine whether it satisfies \(episode < EPISODE\), where \(EPISODE\) is the total number of episodes; if so, then \(episode = episode+1\), and return to step (3); otherwise, the optimization ends, and the optimized reinforcement learning model is obtained.

[0172] According to the above example, data simulation is carried out:

[0173] Assume that in a three-dimensional Cartesian coordinate system, the access point location is \((0, 0, 10)\ m\); the RIS location is \((5, 5, 5)\ m\); within the total number of training time slots \(T = 100\) of the industrial robot, the distance from the RIS changes uniformly from \(5\ m\) to \(15\ m\) with a step size of \(0.1\).

[0174] Assume that in the communication system, the energy transfer efficiency \(\eta=0.7\); the path loss exponent \(\alpha=-1.6\); the energy and signal balance coefficient \(\delta = 0.5\); the path loss \(\kappa=-30\ dB\) at the reference distance \(d' = 1\ m\); the additional loss factor of NLoS The maximum transmit power \(P\) of the access point max \(=80\ W\); the minimum signal quality threshold \(SINR\) min \(=12\ dB\); the proportionality factor \(c\) of the Poisson distribution \(=0.1\).

[0175] In the TD3 deep reinforcement learning, the Actor policy network, Critic comment network 1, and Critic comment network 2 are all composed of fully connected neural networks with three hidden layers, and the AdamPropOptimizer is used as the optimizer. The simulation network environment parameters are the total number of training episodes \(EPISODE = 500\), the total number of training time slots for each episode is \(T = 100\), the random sampling quantity is \(B\) m \(=128\), the interval times \(T\) for updating the Actor policy network and the target network u \(=2\), and the learning rates of the Actor policy network, Critic comment network 1, and Critic comment network 2 are all set to \(0.00001\).

[0176] Figure 3 Show the normalized cumulative reward values corresponding to the total number of training time slots for each episode under the total number of training episodes \(EPISODE = 500\) It can be seen that the reward converges as the training time step increases, and the reward obtained by the TD3 algorithm optimized based on time and space is better than the three benchmark schemes. Since the TD3 algorithm uses two Critic comment networks and updates the Actor policy network, Critic comment network and other policies at an interval of T u = 2, it has a more stable learning process, can learn from the environmental learning and adjust the optimization variables to approximate the optimal solution, and also indicates that the deployment of RIS also plays an important role in improving the wireless energy harvesting communication of RIS-assisted mobile industrial robots.

[0177] Correspondingly, the present application further provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the RIS-assisted mobile industrial robot wireless energy harvesting communication optimization method based on the TD3 algorithm as described above. As Figure 4 shown, it is a hardware structure diagram of any device with data processing capabilities where the RIS-assisted mobile industrial robot wireless energy harvesting communication optimization method based on the TD3 algorithm provided by the embodiment of the present invention is located. In addition to Figure 4 the processors, memory, and network interfaces shown, any device with data processing capabilities where the device in the embodiment is located usually also includes other hardware according to the actual functions of the any device with data processing capabilities, which will not be elaborated here.

[0178] Correspondingly, the present application further provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the RIS-assisted mobile industrial robot wireless energy harvesting communication optimization method based on the TD3 algorithm as described above is implemented. The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both the internal storage unit of any device with data processing capabilities and the external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store the data that has been output or will be output.

[0179] Other embodiments of the present application will be readily apparent to those skilled in the art upon consideration of the specification and practice of the disclosure herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common general knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only illustrative.

[0180] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.

[0181] The above are only the preferred embodiments of the present invention. Although the present invention has been disclosed above in preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make many possible changes and modifications to the technical solution of the present invention, or modify it into equivalent embodiments with equivalent changes, without departing from the scope of the technical solution of the present invention. Therefore, any simple modification, equivalent change, and modification made to the above embodiments according to the technical essence of the present invention without departing from the content of the technical solution of the present invention still fall within the scope of protection of the technical solution of the present invention.

Claims

1. A method for optimizing wireless energy-harvesting communication of RIS-assisted mobile industrial robots based on the TD3 algorithm, characterized in that, The optimization method is based on a RIS-assisted wireless energy-harvesting communication system for mobile industrial robots; the communication system includes an access point, a RIS, and industrial robots; there is only one antenna with D arrays for beamforming on the access point, which can simultaneously transmit energy and signal electromagnetic waves; there are K industrial robots, denoted as each installed with an antenna with a single array; there is only one RIS, which receives the energy and signal electromagnetic waves from the access point and reflects them to the K industrial robots; The RIS-aided wireless energy harvesting communication optimization method for mobile industrial robots based on the TD3 algorithm includes: (1) Modeling the channels between the access point and the RIS, and the channels between the RIS and K robots; (2) Modeling the path losses of the channels between the access point and the RIS, and the channels between the RIS and K robots; (3) Establishing optimization problems regarding energy transfer and information transfer at different time slots t according to the path losses of the channels; (4) Establishing a deep reinforcement learning model according to the channel environment, variables to be optimized, and objectives to be optimized; (5) Optimizing the reinforcement learning model using the TD3 algorithm, and obtaining the solutions to the optimization problems according to the optimized reinforcement learning model, so as to obtain the time ratios and space ratios for deploying the RIS to reflect energy and signals at different time slots t.

2. The RIS-assisted wireless energy-harvesting communication optimization method for mobile industrial robots based on the TD3 algorithm according to claim 1, wherein, (1) Modeling the channels between the access point and the RIS, and the channels between the RIS and K robots includes: (1) Establishing a three-dimensional Cartesian coordinate system for all communication nodes; (2) The RIS is equipped with a total of reflection units arranged in m rows and n columns, and the reflection unit array is denoted as The reflection unit in the i-th row and j-th column is denoted as where i ∈ {1, 2, …, m}, j ∈ {1, 2, …, n}, and the amplitude and phase of each reflection unit can be adjusted; (3) Divide the entire time period into T equal-length time slots, denoted as Among them, each time slot t is divided into two stages, an energy transmission stage and an information transmission stage; in the energy transmission stage, all RIS units are used to transmit energy, and its time accounts for τ(t) in time slot t, where 0 ≤ τ(t) ≤ 1; in the information transmission stage, for industrial robot k, only RIS units are used to transmit information, denoted as while the other part of the RIS units is used to transmit energy; the time of the information transmission stage accounts for 1 - τ(t) in time slot t; denote the ratio of the number of RIS units used to transmit information in the information transmission stage to the total number of RIS units as (4) The beamforming of the access point adopts a linear precoding strategy, and its transmitted signal is expressed as: Among them, and are the electromagnetic waves for energy transmission and information transmission respectively; and are the precoding vectors used by the access point for energy transmission and information transmission respectively; represents a complex vector in the D×1 space; and are the original signals used by the access point for energy transmission and information transmission respectively, both of which follow a circularly symmetric complex Gaussian distribution with a mean of 0 and a variance of 1; the total power p used by the access point for energy transmission is expressed as: Among them, denotes conjugate transpose, and ||·|| denotes the 2-norm of a vector; (5) At each time slot t, the industrial robot The received energy E k (t) is expressed as: Among them, Z k is the baseband equivalent channel from the access point to the RIS unit ; is the baseband equivalent channel from the RIS unit to the industrial robot k; is the diagonal matrix of the reflection coefficient matrix of the RIS unit , where θ l and β l represent the phase and amplitude of the RIS unit l, respectively; is the baseband equivalent channel from the access point to some RIS units ; is the baseband equivalent channel from some RIS units to the industrial robot k; is the diagonal matrix of the reflection coefficient matrix of some RIS units ; η ∈ (0, 1) is the energy transfer efficiency; ω i,j;k (t) is an indicator function indicating whether the RIS unit is used for information transmission, expressed as: In the energy transfer phase with a proportion of τ(t), 0 ≤ τ(t) ≤ 1, in each time slot t, all RIS units are evenly allocated to each industrial robot k; (7) At each time slot t, the industrial robot The received information y k (t) is expressed as: Among them, v k represents the noise of each industrial robot, which follows additive white Gaussian noise with a noise power of ; (8) Considering both the line-of-sight and non-line-of-sight channels simultaneously, and modeling the channel from the RIS to the industrial robot as following a Rayleigh distribution.

3. The RIS-assisted wireless energy-harvesting communication optimization method for mobile industrial robots based on the TD3 algorithm according to claim 1, wherein (1) Modeling the path losses of the channels between the access point and the RIS, and the channels between the RIS and K robots includes: (1) The channel between the access point and the RIS is mainly a line-of-sight (LoS) channel, and the channel path loss between the access point and the RIS unit is modeled as: is modeled as: where (x i,j , y i,j , z i,j ) represents the position of the RIS element in the i-th row and j-th column in the three-dimensional Cartesian coordinate system; the access point is at the origin position of the three-dimensional Cartesian coordinate system; α represents the path loss exponent of the signal path from the access point to the RIS element ; κ represents the path loss of the signal path at the reference distance d'. (2) The access point is assumed to be a point mass in three-dimensional space; (3) The channel between the RIS and the industrial robot includes a line-of-sight channel and a non-line-of-sight (NLoS) channel. The path loss L of the channel between the RIS and the industrial robot RI is modeled as: Among them, respectively represent the probabilities of the LoS channel and the NLoS channel between the RIS and the industrial robot; represents the path loss exponent of the signal path from the RIS to the industrial robot; represents the additional path loss factor under the NLoS channel; d k (t) is the distance from the RIS to each industrial robot k at each time slot t; In the channel between the RIS and the industrial robot, d(t) within each time slot t is considered to be constant; k (t) is considered to be constant; (5) Modeling the LoS channel between the RIS and the industrial robot as: where the variable X represents the number of times the channel between the RIS and the target industrial robot is blocked by other mobile industrial robots per unit time, and X follows a Poisson distribution with parameter λ; the probability distribution function of X is expressed as: where λ is proportional to the unit time, denoted as: where c, c>0 is a proportionality factor used to represent the severity of the blockage; (6) Modeling the NLoS channel between the RIS and the industrial robot as: (7) Considering radio frequency interference, the signal-to-interference-plus-noise ratio (SINR) is used to represent the quality of the signal received by the industrial robot and is defined as: Among them, SINR k (t) is the signal-to-interference-plus-noise ratio of the k-th industrial robot at time slot t.

4. The RIS-assisted wireless energy-harvesting communication optimization method for mobile industrial robots based on the TD3 algorithm according to claim 1, characterized in that, (1) Establishing optimization problems regarding energy transfer and signal transfer at different time slots t according to the path losses of the channels includes: The system objective is to maximize the transmitted energy and communication quality by optimizing the spatial ratio ξ k (t) of the RIS unit for reflecting energy and signals, as well as the time ratio τ(t) / (1 - τ(t)) of solely reflecting energy and reflecting energy and signals simultaneously, subject to the minimum threshold SINR min for communication quality. This optimization problem is formulated as: Among them, is the reward function of the k-th industrial robot at time slot t. δ is the energy and signal balance coefficient, which is used to characterize the importance of communication quality in the entire optimization objective; C1 is the minimum threshold of communication quality, that is, the minimum signal-to-interference-plus-noise ratio constraint SINR min ; C2 is the finite time constraint; C3 is the maximum energy transmission power constraint p max ; C4 is the RIS ratio constraint.

5. The RIS-assisted wireless energy-harvesting communication optimization method for mobile industrial robots based on the TD3 algorithm according to claim 4, wherein (1) Establishing a deep reinforcement learning model according to the channel environment, variables to be optimized, and objectives to be optimized includes: (1) The state space of the deep reinforcement learning model includes: the transmitted signal G of the access point; the distance d k (t) between the RIS and the k-th industrial robot; the baseband equivalent channel h k (t) between the RIS and the k-th industrial robot; (2) The action space of the deep reinforcement learning model includes: the time ratio and the spatial ratio ξ k (t) at time slot t; (3) Model the reward function of the deep reinforcement learning model as the optimization objective described in claim 4 (4) The optimization objective The received dry ratio must not be lower than the minimum threshold SINR min Under the constraint, the reward value of the k-th industrial robot at time slot t is modeled as:

6. The RIS-assisted wireless energy-harvesting communication optimization method for mobile industrial robots based on the TD3 algorithm according to claim 1, wherein In step (5), the TD3 algorithm is used to optimize the reinforcement learning model, and the solution of the optimization problem is obtained according to the optimized reinforcement learning model, so as to obtain the time ratio τ(t) and the spatial ratio ξ(t) of deploying the RIS for reflecting energy and signals at different time slots t, including: k (t), including: (1) Initialize the Actor policy network Critic evaluation network and Critic evaluation network 2 as well as the corresponding Actor target network and Critic target network and Critic target network where s represents the state, a represents the action, and w represents the network parameter vector; (2) Initializing the number of training episodes episode to 0; (3) Initializing the time slot t in the training episode episode to 0; (4) The Actor policy network outputs the action A(t) based on the input state S(t) and obtains the reward value and then transitions to the next state S(t + 1) to obtain the training data set and stores it in the experience replay repository of size B ; (5) Randomly sample B m samples from the experience replay repository to form a data set, and send it to the Actor policy network, Critic evaluation network 1, Critic evaluation network 2, Actor target network, Critic target network 1, and Critic target network 2; among them, the i-th sample is denoted as (6) The objective function is calculated from the minimum Q-value functions of the Critic target network 1 and the Critic target network 2: where y i is the objective function value of the i-th sample; γ is the penalty factor, representing the importance of the future expected Q-value function reward; is the bounded-smoothed target action, obtained by adding bounded noise i+1 to the action A ; (7) The Critic critic network 1 and the Critic critic network 2 are obtained by minimizing the mean square error between the above objective function and the expected Q-value: where j = 1, 2 corresponds to the Critic critic network 1 and the Critic critic network 2 respectively; When the time slot t passes T u times, update the Actor policy network; the parameters of the Actor policy network are updated by maximizing the expected Q value, and the gradient of the above expected Q value is expressed as: When the time slot t passes T u times, update the target network; according to the Actor policy network parameters Critic comment network 1 parameters and Critic comment network 2 parameters respectively and smoothly update the Actor target network parameters Critic target network 1 parameters and Critic target network 2 parameters where ψ ∈ [0, 1] is a smoothing update hyperparameter; (10) Judging whether t < T is satisfied, where T is the total number of time slots; if so, t = t + 1, and return to step (4); otherwise, complete the generation of the RIS strategy for one round of episodes, and continue the training for the next round of episodes. (11) Determine whether the number of episodes episode < EPISODE is satisfied, where EPISODE is the total number of episodes; if so, then episode = episode + 1, and return to step (3); otherwise, the optimization ends, and the optimized reinforcement learning model is obtained.

7. The RIS-assisted wireless energy-harvesting communication optimization method for mobile industrial robots based on the TD3 algorithm according to claim 6, characterized in that, Bounded and smoothed target action By adding bounded noise i+1 to action A to obtain: Among them, A i+1 is the output of the Actor target network, represents a bounded function that is restricted between the lower bound α L and the upper bound α H and then output to That is: if then If then If then It is expressed as: Among them, bounded noise is defined as c and -c are the upper and lower bounds of the bounded function clip for respectively; denotes a normal distribution with a mean of 0 and a variance of respectively.

8. An electronic device, comprising a memory and a processor, characterized in that The memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the RIS-assisted mobile industrial robot wireless energy harvesting communication optimization method based on the TD3 algorithm as described in any one of claims 1-7.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the RIS-assisted mobile industrial robot wireless energy harvesting communication optimization method based on the TD3 algorithm as described in any one of claims 1-7.

10. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, it implements the RIS-assisted mobile industrial robot wireless energy harvesting communication optimization method based on the TD3 algorithm as described in any one of claims 1-7.

Citation Information

Patent Citations

  • IRS-assisted unmanned aerial vehicle communication joint optimization method based on DDPG algorithm

    CN113162679A

  • Intelligent reflection surface assisted vehicle networking safety calculation unloading method, system and equipment and medium

    CN116208619A

  • DDPG-based IRS-assisted cognitive radio system beam forming method

    CN117767987A

  • Unmanned aerial vehicle auxiliary communication sensing trajectory planning method based on reinforcement learning

    CN118778695A

  • Unmanned aerial vehicle secrecy rate optimization method based on evolutionary reinforcement learning

    CN118921654A