Intelligent reflecting surface assisted dual-functional backbone node internet of things resource allocation method

By constructing a resource allocation optimization model and a deep reinforcement learning algorithm, the problems of resource occupation by intelligent reflective surfaces and the concealment of illegal nodes in the Internet of Things are solved. This achieves QoS guarantee for member nodes and normal perception of illegal nodes, and is suitable for wireless communication environments where the line of sight is blocked.

CN118075775BActive Publication Date: 2025-12-05SHANGHAI INST OF MICROSYSTEM & INFORMATION TECH CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410170781.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-06
Publication Date
2025-12-05
Estimated Expiration
2044-02-06

AI Technical Summary

Technical Problem

In the Internet of Things (IoT) of smart cities, intelligent reflective surfaces generate cascaded links in the integrated communication and sensing transmission links, occupying the original transmission link resources, resulting in a decrease in the quality of service (QoS) of member nodes. At the same time, illegal nodes can hide behind obstacles and cannot be detected normally.

Method used

A resource allocation optimization model is constructed. By optimizing the precoding matrix of member nodes, the power covariance matrix of the radar waveform of the perceived target of the backbone node, and the phase shift matrix of the intelligent reflector, combined with deep reinforcement learning algorithms, resource allocation for IoT direct links and cascaded links is realized, ensuring that illegal nodes are properly perceived and maintaining the QoS of member nodes.

Benefits of technology

It enables accurate detection of illegal nodes while ensuring the QoS of member nodes, and adapts to changes in the wireless channel transmission environment, making it particularly suitable for wireless communication scenarios where the line of sight is blocked.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118075775B_ABST
    Figure CN118075775B_ABST
Patent Text Reader

Abstract

The application relates to a kind of intelligent reflecting surface assisted dual-function backbone node internet of things resource allocation methods, comprising: constructing the optimization model of resource allocation, the optimization target is to maximize the beam gain of illegal node, the optimization parameter is the precoding matrix of member node, the power covariance matrix of the sensing target radar waveform of backbone node and the phase shift matrix of intelligent reflecting surface, the constraint condition is that the signal-to-interference-and-noise ratio of each member node is greater than or equal to the preset signal-to-interference-and-noise ratio and the sum of the power of each member node and the power of the sensing target radar waveform of backbone node is less than or equal to the preset power;The optimization model is solved, and the optimal precoding matrix of member node, the power covariance matrix of the sensing target radar waveform of backbone node and the phase shift matrix of intelligent reflecting surface at the current time are obtained.The intelligent reflecting surface assisted dual-function backbone node internet of things resource allocation method of the application can meet the QoS of member node, while ensuring that illegal node can be normally sensed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of Internet of Things (IoT) communication technology, and more specifically to an IoT resource allocation method for dual-function backbone nodes assisted by an intelligent reflective surface. Background Technology

[0002] In future sixth-generation mobile communication systems, the latest advancements in new wireless communication and intelligent application technologies have enabled the rapid development of the Internet of Things (IoT). The IoT, with its ubiquitous sensing and computing capabilities, can interconnect millions of physical objects, forming an indispensable part of today's interconnected world and offering enormous potential for all aspects of modern life. However, the intelligent application technologies in IoT ad hoc networks, such as machine-type communication, interconnected machines, autonomous driving, and extended reality, make it impossible for traditional separation of sensing and communication designs to simultaneously achieve high-speed data transmission and high-precision sensing.

[0003] Integrated sensing technology, developed based on active phased array radar technology, integrates wireless communication and radar sensing by sharing backbone node hardware, effectively addressing the increasingly scarce spectrum resources. This technology is commonly used for communication and target detection in smart city IoT ad hoc network devices. However, unauthorized nodes often hide behind obstacles to actively avoid radar's visual detection area, making them undetectable. This poses a certain uncontrollable risk to the interconnected smart city. Furthermore, the radar signals from the backbone nodes of smart city IoT devices can interfere with member nodes, compromising their QoS. Reconfigurable smart surfaces, due to their ability to endow channels with reconfigurable properties, have received significant attention from academia and industry in recent years. Applying smart reflective surfaces to wireless channel environments where line-of-sight is obstructed allows for the programmatic manipulation of the incident signal phase using passive components. This actively achieves reflection, scattering, and other propagation characteristics of the incident signal, artificially creating a multipath propagation environment that allows unauthorized nodes outside the line-of-sight link to be detected.

[0004] However, intelligent reflective surfaces generate additional cascaded links on top of the integrated communication and sensing transmission links, which occupy the network resources of the original transmission links and affect the quality of service (QoS) of the original member nodes. Summary of the Invention

[0005] The purpose of this invention is to provide a smart reflective surface-assisted dual-function backbone node IoT resource allocation method, which can reasonably allocate resources for direct and cascaded links of the IoT to meet the QoS of member nodes, while ensuring that illegal nodes can be detected normally.

[0006] To achieve the above objectives, this invention provides a dual-function backbone node Internet of Things (IoT) resource allocation method assisted by a smart reflective surface. The IoT includes a backbone node, multiple member nodes, and a smart reflective surface. The backbone node is connected to each member node and the smart reflective surface. The smart reflective surface is located near an obstacle and is connected to each member node and an unauthorized node hidden behind the obstacle. The method includes the following steps:

[0007] S100: Construct an optimization model for resource allocation, wherein the optimization objective of the optimization model is to maximize the beam gain of illegal nodes, the optimization parameters are the precoding matrix of member nodes, the power covariance matrix of the sensing target radar waveform, and the phase shift matrix of the smart reflector, and the constraints include that the signal-to-interference-plus-noise ratio of each member node is greater than or equal to the preset signal-to-interference-plus-noise ratio and the sum of the power of each member node and the power of the sensing target radar waveform of the backbone node is less than or equal to the preset power.

[0008] S200: Solve the optimization model to obtain the optimal precoding matrix of the member nodes, the power covariance matrix of the target radar waveform of the backbone nodes, and the phase shift matrix of the intelligent reflector at the current time, so that the Internet of Things operates according to the optimal precoding matrix of the member nodes, the power covariance matrix of the target radar waveform of the backbone nodes, and the phase shift matrix of the intelligent reflector at the current time.

[0009] Furthermore, the beam gain of the illegal node satisfies the following relationship:

[0010]

[0011] Where ρ is the beam gain of the illegal node, and h r =H B,R Φα(θ r ), H B,R Let Φ be the channel from the backbone node to the smart reflector, Φ be the phase shift matrix of the smart reflector, and w be the channel. k This is the precoding vector corresponding to the k-th member node. R represents the set of member nodes in the Internet of Things (IoT). r The power covariance matrix of the radar waveform for sensing the target.

[0012] Furthermore, the power covariance matrix of the target radar waveform of the backbone node satisfies the following relationship:

[0013]

[0014] Where, x r It is the target sensing radar waveform transmitted by the dual-function backbone node. For x r The conjugate matrix, For expectations;

[0015] The precoding matrices of member nodes satisfy the following relationship:

[0016] W = [w1, w2, ... w K ]

[0017] Where W is the precoding matrix of the member nodes, and K is the total number of member nodes.

[0018] Furthermore, the signal-to-interference-plus-noise ratio γ of the k-th member node k The following relationship must be satisfied:

[0019]

[0020] Among them, H k For cascaded channels of member nodes, It is the thermal noise power of the kth member node.

[0021] Furthermore, the constraints satisfy the following relationship:

[0022] γ k ≥γ req

[0023]

[0024] Where, γ req To preset the signal-to-interference-plus-noise ratio, Tr(R) r ) is a matrix R r traces, This is the preset power.

[0025] Furthermore, step S200 specifically includes:

[0026] S210: Using the backbone node as the agent, the optimization parameters as the action of the agent, the signal-to-interference-plus-noise ratio of each member node as the state, and the optimization objective as the reward, the optimization model is transformed into a Markov decision process.

[0027] S220: Deploy the pre-learned optimal policy into the agent;

[0028] S230: The agent acquires the current state and outputs the optimal action at the current moment based on the current state and the optimal strategy. The optimal action includes the precoding matrix of the optimal member node, the power covariance matrix of the target radar waveform of the backbone node, and the phase shift matrix of the intelligent reflector.

[0029] Furthermore, in step S220, the learning method for the optimal policy specifically includes:

[0030] S221: Initialize two online neural networks, two target neural networks, one policy neural network, and one temperature neural network. The two online neural networks and the two target networks take state and action as input and the value of action as output. The policy neural network takes state as input and action as output. The temperature neural network takes state and policy as input and temperature parameter as output. Initialize the learning rate and soft update factor, and set the initial number of iterations I = 1.

[0031] S222: Let t = 1, initialize state s1. If I = 1, let s1 = 0; otherwise, let s1 = s tmax ; where t max The maximum time value;

[0032] S223: The agent randomly executes action a t It interacts with the environment to obtain the state s t+1 and reward r t+1 , will a t s t s t+1 r t+1 Saved in the experience pool;

[0033] S224: Randomly select a batch of tuples from the experience pool, and use each tuple to train two online neural networks in order to update the parameters θ1 and θ2 of the two online neural networks;

[0034] S225: Train the policy neural network using each tuple to update the parameters of the policy neural network.

[0035] S226: Train the temperature neural network using each tuple to update the parameters α of the temperature neural network;

[0036] S227: Update the parameters of the two target neural networks

[0037] S228: Determine if time t is less than the maximum time value t max If yes, increment time t by 1 (i.e., t = t + 1), and then return to step S223; otherwise, proceed to step S229.

[0038] S229: Determine if the iteration count I is less than the maximum iteration count I. max If yes, increment the iteration count by 1 and return to step S222; otherwise, learning is complete, and the learned policy neural network is obtained as the optimal policy.

[0039] Furthermore, action a t It can be any action in the action space, which is defined by the constraints.

[0040] Furthermore, in step S224, the two online neural networks are trained with the goal of minimizing the difference between the estimated value and the target value of the online neural network.

[0041] Further, in step S227, the updated parameters of the two target neural networks are obtained based on the parameters of the two online neural networks, the current parameters of the two target neural networks, and the soft update factor.

[0042] The present invention provides a smart reflector-assisted dual-function backbone node IoT resource allocation method. By solving an optimization model for resource allocation, the optimal precoding matrix of member nodes, the power covariance matrix of the target radar waveform of the backbone node, and the phase shift matrix of the smart reflector can be obtained at the current moment. This achieves optimal resource allocation, ensuring the QoS of each member node while guaranteeing that illegal nodes are detected normally. In the optimization model, the smart reflector is introduced and the target beam gain is used as the optimization objective, thus accurately detecting illegal nodes and ensuring the QoS of each member node. The optimization model is solved based on a deep reinforcement learning algorithm, which can adapt to the ever-changing wireless channel transmission environment in the IoT, and is particularly suitable for deployment in scenarios where visual links are blocked and wireless communication environments are variable. Attached Figure Description

[0043] Figure 1 This is a schematic diagram of the structure of the Internet of Things according to an embodiment of the present invention;

[0044] Figure 2 A flowchart of an IoT resource allocation method for a dual-function backbone node assisted by a smart reflective surface according to an embodiment of the present invention;

[0045] Figure 3 This is a schematic diagram of a Markov decision process according to an embodiment of the present invention. Detailed Implementation

[0046] The preferred embodiments of the present invention are given below with reference to the accompanying drawings and described in detail.

[0047] like Figure 1As shown, the Internet of Things (IoT) of this embodiment includes a backbone node 10, K member nodes 20, and a smart reflector (RIS) 30. The backbone node 10 includes N antennas, enabling simultaneous communication and sensing. The member nodes 20 are devices used to collect data from the physical world, such as sensors, actuators, cameras, GPS locators, etc. The smart reflector 30 is positioned near an obstacle 40 and is connected to each member node 20 and an illegal node 50 hidden behind the obstacle 40. The backbone node 10 is connected to each member node 20 and the smart reflector 30. The smart reflector 30 contains L reflective elements, which together form a phase shift matrix Φ, satisfying Φ = diag(φ1, φ2, ..., φ). L ), φ L This represents the phase shift (i.e., phase offset) of the Lth reflecting element. A direct channel, denoted as H, is formed between the backbone node 10 and each member node 20. B,k Where k (k = 1, 2, ..., K) represents the k-th member node, and B refers to backbone node 10. The channel from backbone node 10 to smart reflector 30 can be H B,R This indicates that R represents the smart reflector. The channel from the smart reflector 30 to each member node 20 is available via h. R,k The cascaded channel of member node 20 (i.e., the channel connected to backbone node 10 via smart reflector 30) is: H k =H B,k +H B,R Φh R,k .

[0048] The transmission waveform of backbone node 10, which includes communication and sensing, can be represented as follows:

[0049]

[0050] Where, x k w represents the transmission symbol of the k-th member node 20. k This is the precoding vector corresponding to the k-th member node 20. The precoding vectors of all member nodes 20 constitute the precoding matrix, and the precoding matrix W = [w1, w2, ... w2]. K ], x represents the set of member nodes in the Internet of Things (IoT). r It is the target sensing radar waveform sent by the dual-function backbone node.

[0051] The received signal of the kth member node 20 is denoted as:

[0052]

[0053] Where, n k This represents the thermal noise received by the k-th member node.

[0054] like Figure 2 As shown, the IoT resource allocation method for a dual-function backbone node assisted by a smart reflective surface according to an embodiment of the present invention includes the following steps:

[0055] S100: Construct an optimization model for resource allocation; the optimization objective of the model is to maximize the beam gain ρ of the illegal node 50, and the optimization parameters are the precoding matrix W of member node 20 and the power covariance matrix R of the radar waveform of the sensing target of backbone node 10. r The phase shift matrix Φ of the intelligent reflector 30 is subject to the following constraints: the signal-to-interference-plus-noise ratio (SIR) of each member node 20 is greater than or equal to the preset SIR, and the sum of the power of each member node 20 and the power of the target radar waveform of the backbone node 10 is less than or equal to the preset power.

[0056] The beam gain ρ of illegal node 50 satisfies the following relationship:

[0057]

[0058] Among them, h r =H B,R Φα(θ r ), α(θ) r ) is the guiding vector of illegal node 50, θ r It is the pitch angle of illegal node 50.

[0059] The power covariance matrix R of the target sensing radar waveform of backbone node 10 r The following relationship must be satisfied:

[0060]

[0061] in, For x r The conjugate matrix, As expected.

[0062] The signal-to-interference-plus-noise ratio γ of the k-th member node k for:

[0063]

[0064] in, It is the thermal noise power of the kth member node.

[0065] Therefore, the optimization model can be expressed as:

[0066]

[0067] stγ k ≥γ req

[0068]

[0069] Where, γ req To preset the signal-to-interference-plus-noise ratio, For preset power, The sum of the power of each member node 20, Tr(R) r ) is the trace of the power covariance matrix, which is the power of the radar waveform of the perceived target at backbone node 10.

[0070] S200: Solve the optimization model to obtain the optimal precoding matrix W of member node 20 and the power covariance matrix R of the target radar waveform of backbone node 10 at the current time. r The phase shift matrix Φ of the intelligent reflector 30 is used to enable the Internet of Things to operate according to the optimal precoding matrix W of the member node 20 and the power covariance matrix R of the target radar waveform of the backbone node 10 at the current moment. r It operates in conjunction with the phase shift matrix Φ of the intelligent reflector 30 to achieve optimal resource allocation.

[0071] The optimized W and R r And Φ, the backbone node 10 can be configured with the optimized W to configure the precoding method of each member node 20, and with the optimized R r The power sensing illegal node 50 is used, and the reflection phase shift of the smart reflector 30 is configured with the optimized Φ, so that the Internet of Things is in the optimal W and R at the current moment. r The system operates in conjunction with Φ to achieve optimal resource allocation, ensuring the QoS (signal-to-interference-plus-noise ratio) of each member node 20 while also guaranteeing that illegal nodes 50 are properly detected.

[0072] In some embodiments, the above optimization model can be solved based on a deep reinforcement learning algorithm (e.g., the SAC algorithm). Specifically, step S200 includes:

[0073] S210: Using backbone node 10 as the agent, optimization parameters as the agent's actions, the signal-to-interference-plus-noise ratio of each member node as the state, and the optimization objective as the reward, the optimization model is transformed into a Markov decision process.

[0074] S220: Deploy the pre-learned optimal policy in the agent;

[0075] S230: The agent acquires the current state and outputs the optimal action for the current moment based on the current state and the optimal policy, namely the optimal precoding matrix W of member node 20 and the power covariance matrix R of the radar waveform of the perceived target of backbone node 10. r And the phase shift matrix Φ of the intelligent reflective surface 30.

[0076] like Figure 3 As shown, in the Markov decision-making process, the agent (i.e., backbone node 10) interacts with the environment (the rest of the Internet of Things excluding backbone node 10). Specifically, the agent can obtain the state of the environment, then map the state to an action according to the policy, and implement the action. After the action is implemented, the state of the environment will change, and the environment will generate a reward and return it to the agent. Then the agent outputs a new action according to the new state and repeats the above process in a loop.

[0077] The agent's action at time t can be represented as:

[0078]

[0079] Here, the subscript t represents time t, and the vec operation transforms the matrix structure into a row vector.

[0080] The state of the environment at time t is the signal-to-interference-plus-noise ratio (SINR) of each member node at time t, which can be denoted as: γ k,t Action a t The corresponding reward is ρ t .

[0081] In some embodiments, the optimal policy can be learned based on the SAC algorithm. The method for learning the optimal policy includes the following steps:

[0082] S221: Initialize two online neural networks, two target neural networks, one policy neural network, and one temperature neural network. The two online neural networks and the two target networks take the state and action as input and the value of the action as output. The policy neural network takes the state as input and the action as output. The temperature neural network takes the state and policy as input and the temperature parameter as output. Initialize the learning rate and soft update factor, and set the number of iterations I = 1.

[0083] Suppose the parameters of the two online neural networks are θ1 and θ2, and the parameters of the two target neural networks are... The parameters of the policy neural network are The temperature neural network has parameter α. During initialization, θ1, θ2, ... Both α and α are set to 1.

[0084] Assuming the learning rate is v and the soft update factor is η, during initialization, v = 0.0001 and ρ = 0.005, and the iteration number I is used to control the number of iterations during training.

[0085] S222: Let t = 1, initialize state s1. If I = 1, let s1 = 0; otherwise, let...

[0086] Among them, t max This is the maximum time value, which can be set as needed.

[0087] S223: The agent randomly executes action a t It interacts with the environment to obtain the state s t+1 and reward r t+1 , will a t s t s t+1 r t+1 Saved in the experience pool.

[0088] Action a t It can be any action in the action space A, where the action space A is limited by constraints.

[0089] S224: Randomly select a batch of tuples from the experience pool, and use each tuple to train two online neural networks in order to update the parameters θ1 and θ2 of the two online neural networks.

[0090] This batch of tuples consists of multiple tuples, each containing a. t s t s t+1 r t+1 When training two online neural networks, the target is the mean squared error between the estimated value and the target value of the online neural network. Specifically, for each tuple, a can be... t and s t Input two online neural networks In the middle, the estimated values ​​were obtained respectively. and Then it is fed into a two-target neural network. In the middle, the target estimates were obtained respectively. and Then, based on the mean squared error between the estimated values ​​of the online neural network and the target estimated values ​​of the target neural network, the loss function of the online neural network is constructed, and the parameters θ1 and θ2 are updated according to the following formulas:

[0091] Where i = 1, 2

[0092] Among them, J Q (θ i Let ) be the loss function, which can be expressed as:

[0093]

[0094]

[0095] Where ω is the discount factor and α is the temperature factor (i.e., the parameters of the temperature neural network). For a policy neural network, i = 1, 2.

[0096] S225: Train the policy neural network using each tuple to update the parameters of the policy neural network.

[0097] Policy Neural Network The loss can be expressed as

[0098]

[0099] Where i = 1, 2.

[0100] Its update formula is:

[0101]

[0102] S226: Train the temperature neural network using each tuple to update the parameters α of the temperature neural network.

[0103] The loss function J(α) of the temperature neural network is:

[0104]

[0105] Where H is the equivalent hyperparameter of the target entropy.

[0106] The update formula for parameter α is:

[0107]

[0108] S227: Update the parameters of the two target neural networks

[0109] In some embodiments, a soft update method can be used to update the parameters of the target neural network. This involves obtaining the updated parameters of the target neural network based on a soft update factor, the current parameter values ​​of the target neural network, and the parameter values ​​of the online neural network. Specifically, the update method is as follows:

[0110]

[0111] S228: Determine if time t is less than the maximum time value t max If yes, increment time t by 1 (i.e., t = t + 1), and then return to step S223; otherwise, proceed to step S229.

[0112] S229: Determine if the iteration count I is less than the maximum iteration count I. max If yes, increment the iteration count by 1 and return to step S222; otherwise, learning is complete, and the learned policy neural network is obtained as the optimal policy.

[0113] After the optimal strategy is trained, it can be deployed on backbone node 10. Backbone node 10 outputs the optimal strategy configuration based on the environment of each member node at the current time. The optimal strategy includes the precoding matrix W of member node 20 and the power covariance matrix R of the radar waveform of the perceived target of backbone node 10. r The phase shift matrix Φ of the intelligent reflector 30 is used to enable the Internet of Things to operate according to the optimal precoding matrix W of the member node 20 at the current time, and the power covariance matrix R of the target radar waveform of the backbone node 10. r It operates in conjunction with the phase shift matrix φ of the intelligent reflector 30 to achieve optimal resource allocation.

[0114] In some embodiments, the policy and learning algorithm can be directly deployed on the backbone node 10 so that the policy can be trained by the learning algorithm to obtain the optimal policy. Then, the backbone node 10 can implement resource allocation for the Internet of Things based on the optimal policy. When it is necessary to update the optimal policy, step S220 above can be re-executed to retrain the policy neural network.

[0115] The intelligent reflector-assisted dual-function backbone node IoT resource allocation method of this invention, by solving the resource allocation optimization model, can obtain the optimal precoding matrix W of member node 20 and the power covariance matrix R of the sensing target radar waveform of backbone node 10 at the current moment. r The phase shift matrix φ of the intelligent reflector 30 is used to achieve optimal resource allocation, ensuring the QoS of each member node 20 while also ensuring that the illegal node 50 is properly detected. In the optimization model, the intelligent reflector is introduced and the target beam gain is used as the optimization objective, so the illegal node 50 can be accurately detected and the QoS of each member node 20 can be guaranteed. The optimization model is solved based on the deep reinforcement learning algorithm, which can adapt to the ever-changing wireless channel transmission environment in the Internet of Things, and is particularly suitable for deployment in scenarios where the visual link is blocked and the wireless communication environment is changeable.

[0116] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of the invention. Various variations can be made to the above embodiments of the present invention. That is, all simple and equivalent changes and modifications made based on the claims and description of this invention fall within the protection scope of the claims of this patent. All aspects not described in detail in this invention are conventional technical content.

Claims

1. A method for smart reflector assisted dual-functional backbone node Internet of Things resource allocation, characterized in that, The Internet of Things comprises a backbone node, a plurality of member nodes and a smart reflecting surface, the backbone node is connected with each member node and the smart reflecting surface respectively, the smart reflecting surface is arranged near an obstacle, the smart reflecting surface is connected with each member node and an illegal node hidden behind the obstacle respectively, and the method comprises the following steps: S100: An optimization model of resource allocation is constructed, wherein an optimization target of the optimization model is to maximize the beam gain of the illegal node, optimization parameters are a precoding matrix of the member node, a power covariance matrix of a sensing target radar waveform and a phase shift matrix of the smart reflecting surface, and constraint conditions comprise that a signal-to-interference-and-noise ratio (SINR) of each member node is greater than or equal to a preset SINR, and a sum of a power of each member node and a power of the sensing target radar waveform of the backbone node is less than or equal to a preset power; S200: The optimization model is solved to obtain an optimal precoding matrix of the member node, an optimal power covariance matrix of the sensing target radar waveform of the backbone node and an optimal phase shift matrix of the smart reflecting surface at a current time, so that the Internet of Things operates according to the optimal precoding matrix of the member node, the optimal power covariance matrix of the sensing target radar waveform of the backbone node and the optimal phase shift matrix of the smart reflecting surface at the current time; The beam gain of the illegal node satisfies the following relationship: , wherein, beam gain of the illegal node, , is the conjugate matrix of , is the steering vector of the illegal node, is the elevation angle of the illegal node, is the channel from the backbone node to the smart reflector, is the phase shift matrix of the smart reflector, is the precoding vector corresponding to the kth member node, is the conjugate matrix of , denotes the set of member nodes in the Internet of Things, is the power covariance matrix of the sensing target radar waveform. 2.The smart reflector assisted dual-functional backbone node internet of things resource allocation method according to claim 1, characterized in that, The power covariance matrix of the sensing target radar waveform of the backbone node satisfies the following relationship: , wherein, is a sensing target radar waveform transmitted by the dual-function backbone node, is is a conjugate matrix of is an expectation; The precoding matrix of the member node satisfies the following relationship: , wherein is the precoding matrix of the member node, and K is the total number of member nodes. 3.The smart reflector assisted dual-functional backbone node internet of things resource allocation method according to claim 2, characterized in that, signal-to-interference-plus-noise ratio of the kth member node satisfies the following relation: , wherein, is the concatenated channel for the member nodes, is the thermal noise power of the kth member node. 4.The smart reflector assisted dual-functional backbone node internet of things resource allocation method according to claim 3, characterized in that, The constraint condition satisfies the following relationship: , , wherein is a predetermined signal-to-noise ratio, is a matrix is a trace of the matrix is a predetermined power.

5. The smart reflector assisted dual function backbone node internet of things resource allocation method according to claim 1, characterized in that, Step S200 specifically comprises: S210: The backbone node is taken as an agent, the optimization parameters are taken as actions of the agent, the SINR of each member node is taken as a state, and the optimization target is taken as a reward, so that the optimization model is converted into a Markov decision process; S220: An optimal strategy learned in advance is deployed in the agent; S230: The agent obtains a state at a current time, and outputs an optimal action at the current time according to the state at the current time and the optimal strategy, wherein the optimal action comprises the optimal precoding matrix of the member node, the optimal power covariance matrix of the sensing target radar waveform of the backbone node and the optimal phase shift matrix of the smart reflecting surface.

6. The smart reflector assisted dual function backbone node internet of things resource allocation method according to claim 5, characterized in that, In step S220, a learning method of the optimal strategy specifically comprises: S221: Two online neural networks, two target neural networks, a strategy neural network and a temperature neural network are initialized, wherein the two online neural networks and the two target networks take a state and an action as inputs and take a value of the action as an output; the strategy neural network takes the state as an input and takes the action as an output; the temperature neural network takes the state and the strategy as inputs and takes a temperature parameter as an output; a learning rate and a soft update factor are initialized, and an initial iteration number I is set as 1; S222: Let t = 1, initialize state , if , let , else let ; wherein t max is the maximum time value; S223: the agent randomly performs an action , and interacts with the environment, obtaining a state and a reward , which are saved in the experience pool , , , ​ S224: randomly draw a batch of tuples from the experience pool, train the two online neural networks using each tuple to update the parameters of the two online neural networks 、 ; S225: training the policy neural network with each tuple to update parameters of the policy neural network ; S226: training the temperature neural network with each tuple to update parameters of the temperature neural network ; S227: update parameters of the two target neural networks , ; S228: judge whether the time t is less than the maximum time value t max , if yes, let the time t add 1 (i.e. t=t+1), and then return to step S223; otherwise, enter step S229; S229: Determine whether the iteration number I is less than the maximum iteration number I max If yes, then let the iteration number be incremented by 1, and then return to step S222; otherwise, the learning is completed, and a learned policy neural network is obtained as an optimal policy.

7. The smart reflector assisted dual function backbone node internet of things resource allocation method according to claim 6, characterized in that, Action Any action in an action space defined by the constraints can be. 8.The smart reflector assisted dual-functional backbone node internet of things resource allocation method according to claim 6, wherein, In step S224, the two online neural networks are trained by minimizing a difference between an estimated value and a target value of the online neural network. 9.The smart reflector assisted dual-functional backbone node internet of things resource allocation method according to claim 6, wherein, In step S227, updated parameters of the two target neural networks are obtained according to parameters of the two online neural networks, current parameters of the two target neural networks and the soft update factor.

Citation Information

Patent Citations

  • Multi-user communication system physical layer control method under non-ideal hardware condition

    CN115882911A