An Online Throughput Optimization Method for Wireless Communication Systems Based on Smart Reflectors

By optimizing the energy management of the intelligent reflector through online time switching strategy and Markov decision process, the energy consumption problem of the intelligent reflector in wireless communication system is solved, achieving battery energy neutrality and maximizing throughput, thereby improving the performance of the communication system.

CN119562305BActive Publication Date: 2025-10-28FOSHAN UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411741604.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-29
Publication Date
2025-10-28
Estimated Expiration
2044-11-29

AI Technical Summary

Technical Problem

The energy consumption problem of existing smart reflectors in wireless communication systems has not been effectively solved, especially in battery-assisted IRS networks, where battery energy reserves are limited and it is difficult to adapt to dynamic channel changes in real time, affecting communication stability and throughput.

Method used

By constructing a battery-powered intelligent reflector-assisted wireless communication system and adopting an online time-switching strategy, the system is divided into energy harvesting and signal reflection stages. By combining Markov decision processes and strategy iteration algorithms, the time-switching strategy is optimized to achieve battery energy neutrality and maximize long-term average throughput.

Benefits of technology

In dynamic environments, it adaptively maintains the remaining battery power, enhances the reliability and efficiency of the communication system, ensures the continuity and stability of communication, and maximizes the long-term average throughput of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119562305B_ABST
    Figure CN119562305B_ABST
Patent Text Reader

Abstract

This invention provides an online throughput optimization method for a wireless communication system based on a smart reflector. The method includes constructing a battery-powered smart reflector-assisted wireless communication system; dividing the system into an energy harvesting phase and a signal reflection phase based on an online time-switching strategy; and establishing a long-term average throughput maximization model with the goal of maximizing the long-term average throughput of the system while achieving battery energy neutrality. The smart reflector-assisted wireless communication system executes data transmission according to the optimal time-switching strategy and updates the remaining battery power status in real time. Compared with existing technologies, this invention, upon the arrival of actual channel conditions, executes the system's decisions and updates the remaining battery power status, ensuring the continuity and stability of communication in dynamically changing wireless network environments and achieving battery energy neutrality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of wireless communication technology, and in particular to an online throughput optimization method for wireless communication systems based on intelligent reflective surfaces. Background Technology

[0002] With the rapid development of wireless networks, Intelligent Reflecting Surfaces (IRS) have been extensively studied for improving signal coverage and communication efficiency. IRSs offer flexible deployment, easily installed on building facades, billboards, or other structures, allowing for seamless integration into existing wireless communication systems. However, most existing IRS optimization strategies focus on improving signal processing or propagation models, neglecting the IRS's energy consumption. In battery-powered IRS-assisted wireless communication networks, the IRS is powered by a small battery, but the battery's energy reserves are limited and require periodic replacement or charging as energy is depleted, which is extremely inconvenient in resource-constrained or difficult-to-manage environments. While energy harvesting technologies (such as radio frequency (RF) energy harvesting) provide a viable solution for continuously powering the IRS, the dynamic nature of channels in wireless environments means that traditional channel estimation and energy management strategies often cannot adapt to environmental changes in real time, reducing energy utilization efficiency and affecting signal quality stability. In particular, the states of the incident and reflected channels on both sides of the IRS are constantly changing, making it difficult for the system to obtain channel state information in a timely and accurate manner. This increases the decision-making complexity in achieving energy-neutral operation of the IRS and maximizing long-term average throughput.

[0003] Therefore, there is an urgent need for a new online throughput optimization method for wireless communication systems based on intelligent reflective surfaces to solve the above-mentioned technical problems. Summary of the Invention

[0004] This invention proposes an online throughput optimization method for wireless communication systems based on intelligent reflectors. The aim is to enable the intelligent reflector to adaptively maintain a certain amount of remaining power under different dynamic environments through an optimal time switching strategy, thereby ensuring the continuity and stability of communication in dynamically changing wireless network environments.

[0005] This invention proposes an online throughput optimization method for wireless communication systems based on intelligent reflectors, comprising the following steps:

[0006] S1. Construct a battery-powered intelligent reflective surface-assisted wireless communication system;

[0007] S2. Based on the online time switching strategy, the intelligent reflector-assisted wireless communication system is divided into an energy harvesting stage and a signal reflection stage, and a long-term average throughput maximization model is established with the goal of maximizing the long-term average throughput of the intelligent reflector-assisted wireless communication system and achieving battery energy neutrality.

[0008] S3. Solve the long-term average throughput maximization model based on Markov decision process, strategy iteration algorithm and objective reward function to obtain the optimal time switching strategy;

[0009] S4. The intelligent reflective surface-assisted wireless communication system performs data transmission according to the optimal time switching strategy and updates the remaining power status of the battery in the intelligent reflective surface-assisted wireless communication system in real time.

[0010] Preferably, the intelligent reflector-assisted wireless communication system includes a base station, an intelligent reflector equipped with a rechargeable battery, and multiple single-antenna sensor nodes.

[0011] Preferably, the long-term average throughput maximization model satisfies the following relationship:

[0012]

[0013] st0≤E c (n)≤E a (n+τ(n))

[0014] 0≤E a (n+1)≤E max

[0015] 0≤τ(n)≤1

[0016] E c (n)≤N r P u ;

[0017] Among them, P u N represents the power consumption of a single reflective element in the intelligent reflective surface. r E represents the number of reflective elements in the intelligent reflective surface. c τ(n) represents the energy consumption of the intelligent reflector in time slot n, and τ(n) represents the time allocation of the energy harvesting phase of the intelligent reflector in time slot n. a (n+τ(n)) represents the remaining battery energy of the smart reflector after the energy harvesting phase in time slot n, E max E represents the maximum battery capacity of the intelligent reflective surface. a (n) represents the battery energy state starting at time slot n, Jπ represents the long-term average throughput, C(n) represents the throughput at time slot n under the execution strategy, and Ea (n+1) represents the battery energy state at time slot n+1.

[0018] Preferably, step S3 includes the following sub-steps:

[0019] S31. Establish a system state model based on Markov decision process, and initialize the channel state transition matrix and battery power status.

[0020] S32. Predict the future channel state based on the system state model, and determine the battery energy neutrality boundary based on historical data from the sliding window.

[0021] S33. Iteratively optimize the system state model according to the strategy iteration algorithm;

[0022] S34. Determine whether the current solution is the optimal solution; if yes, output the current solution as the optimal time switching strategy; if no, return to step S33.

[0023] Preferably, step S31 includes the following sub-steps:

[0024] S311. Establish a state space based on the channel state of the first transmission channel, the channel state of the second transmission channel, and the remaining battery power state of the smart reflective surface.

[0025] S312. Establish an action space based on the time ratio of the battery in the absorption state and the reflection state of the intelligent reflective surface;

[0026] S313. Establish the state transition function;

[0027] S314. Establish the target reward function to obtain the system state model.

[0028] Preferably, the target reward function satisfies the following relationship:

[0029]

[0030] Among them, R e (t) represents the target reward function, E a (t) represents the battery energy state in time slot t, E L (t) and E U (t) represents the lower limit and upper limit of battery energy neutrality, respectively, and π represents a constant.

[0031] Preferably, the optimal time switching strategy satisfies the following relationship:

[0032]

[0033] Where, τ *(t) represents the optimal time switching strategy. Let γ represent the average reward of the s′th state of the system state model under policy τ, γ represent the discount factor, and a(t)∈A represent the action policy chosen in action space A during time slot t. e (t) represents the target reward function, T represents the state transition function, and a represents the action taken.

[0034] Compared to existing technologies, this invention utilizes a sliding window to update historical data in real time within each time slot. This historical data includes the incident and reflected channel states of past time slots, the difference between energy harvesting and reflected energy consumption of the intelligent reflector, and the battery energy status. Using the energy difference data and battery energy status, the system obtains an initial energy boundary. An energy boundary adjustment factor is calculated based on the channel states of past time slots, and this factor is used to adjust the initial boundary to obtain an energy-neutral target boundary. An energy optimization method based on Markov processes and online time-switching strategies is employed. According to the designed target reward function, when the overall channel conditions on both sides of the intelligent reflector are poor, the system stores energy through an online time-switching strategy. When the overall channel conditions of the intelligent reflector are good, efficient data transmission is implemented. An optimal time-switching strategy based on predicted states is obtained through strategy iteration algorithm optimization. Upon the arrival of the actual channel environment, the system executes its decisions and updates the remaining battery power status. This allows the intelligent reflector to adaptively maintain a certain amount of remaining power under different dynamic environments, significantly enhancing the reliability and efficiency of the communication system, while ensuring the continuity and stability of communication in dynamically changing wireless network environments. This enables the intelligent reflective surface to operate in an energy-neutral manner and maximizes the long-term average throughput of the system, thereby greatly improving the overall performance of the wireless communication system. Attached Figure Description

[0035] The present invention will now be described in detail with reference to the accompanying drawings. The above and other aspects of the present invention will become clearer and more readily understood through the detailed description following the accompanying drawings. In the drawings:

[0036] Figure 1 This is a schematic diagram of a battery-powered smart reflector-assisted wireless communication system structure, which is based on the online throughput optimization method for a smart reflector-based wireless communication system provided in this embodiment of the invention.

[0037] Figure 2 This is a flowchart of the online throughput optimization method for a wireless communication system based on a smart reflector provided in an embodiment of the present invention;

[0038] Figure 3 This is a logical schematic diagram of the online time switching strategy of the online throughput optimization method for a wireless communication system based on a smart reflector provided in an embodiment of the present invention;

[0039] Figure 4 This is a schematic diagram of the target reward function result when the power constraint of the online throughput optimization method for a wireless communication system based on a smart reflector provided in this embodiment of the invention is [0.3, 0.7].

[0040] Figure 5 This is a schematic diagram showing the change in channel environment gain on both sides of the intelligent reflector in the online throughput optimization method for a wireless communication system based on an intelligent reflector provided in an embodiment of the present invention.

[0041] Figure 6 This is a schematic diagram comparing the long-term average throughput of an online energy management optimization method for smart reflectors and an offline, known channel state of a smart reflector-assisted wireless communication system, as provided in the embodiments of the present invention.

[0042] Figure 7 This is a schematic diagram of the neutral state of the remaining battery power of the intelligent reflector in the online throughput optimization method for a wireless communication system based on an intelligent reflector provided in an embodiment of the present invention. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0044] Please refer to Figures 1-7 This invention proposes an online throughput optimization method for wireless communication systems based on intelligent reflectors, comprising the following steps:

[0045] S1. Construct a battery-powered intelligent reflective surface-assisted wireless communication system;

[0046] In this embodiment of the invention, the intelligent reflector-assisted wireless communication system includes a base station, an intelligent reflector equipped with a rechargeable battery, and multiple single-antenna sensor nodes.

[0047] Specifically, the battery-powered smart reflector-assisted wireless communication system includes N t Base station (BS), N with one transmitting antenna r It consists of a passive reflective element fixed in a position, and K single-antenna sensor nodes. For example... Figure 1 As shown.

[0048] The wireless transmission channel from the BS to the IRS is defined as denoted as And the wireless transmission channel from the IRS to the end user is denoted as The channel gains G and F on both sides of the IRS k The following formulas are given respectively.

[0049]

[0050] in and ε represents the large-scale path loss from the BS to the IRS plane and from the IRS plane to the sensor node, respectively. g and ε f These represent the small-scale attenuation of the channel gain between the BS and IRS, and between the IRS and the sensor node, respectively, reflecting the randomness and instability of the channel. It is assumed that the wireless channel transmission is block fading, meaning the channel fading remains constant within a time slot t, but the channel gain is time-varying throughout the entire time slot sequence.

[0051] The wireless channel gains on both sides of the IRS are modeled as a stationary, time-homogeneous, and irreducible Markov channel model, with a state space of C = {C1, C2, ..., C}. M}, where each state C i Let represent different channel fading gains. Then the state transition of the Markov process can be expressed as: This indicates that the channel gain state at time t+1 depends only on the state at time t and is independent of previous states. Furthermore, this invention assumes that the historical channel state information (HCSI) on both sides of the IRS is known. Within a time slot, the signal received by the relay IRS in the system can be represented by the following vector form.

[0052]

[0053] Where P t It is the base station transmit power, w∈C M×1 The energy carrier signal vector transmitted from the BS, n irs This represents the noise vector at the IRS. The energy received by the IRS is represented as...

[0054] P irs (t)=ηP t ||G(t)w|| 2 (4)

[0055] Where η (0 < η < 1) represents the energy conversion efficiency, which depends on the rectification process and the energy harvesting circuit. The effect of noise power on the received energy is ignored here.

[0056] S2. Based on the online time switching strategy, the intelligent reflector-assisted wireless communication system is divided into an energy harvesting stage and a signal reflection stage, and a long-term average throughput maximization model is established with the goal of maximizing the long-term average throughput of the intelligent reflector-assisted wireless communication system and achieving battery energy neutrality.

[0057] In this embodiment of the invention, after the signal energy is transmitted to the IRS, the IRS divides the time slot t into two stages, energy acquisition (EH) and signal reflection (SR), by using the time factor τ∈[0,1].

[0058] A time slot t is divided into two phases: energy harvesting (EH) and signal reflection (SR). The battery capacity (E) is defined. max The remaining battery charge in time slot t is E. a (t), is known only at the end of time slot t-1. The IRS adjusts the duration of these two phases within a complete time slot t using a time factor τ∈[0,1]. During time τ, the IRS adjusts the amplitude β of the reflecting unit. n A value of 0 allows the IRS to collect energy and store it in the battery. In the optimal beam ||w|| 2 When = 1, during the EH phase, the energy harvesting of the IRS is

[0059] e(t)=τ(t)ηP t L0MNε g 2 (t) (5)

[0060] At this time, the energy of the IRS battery is E. a (t+τ(t))=E a (t)+e(t). During the time interval (1-τ), the IRS sets its reflection amplitude to 1, completely reflecting the energy signal from the BS back to the user. During this process, the IRS no longer absorbs energy, and its power consumption is provided by the energy stored in the battery. Therefore, the signal received at the k-th sensor user is represented as:

[0061] y k (t)=F k (t)Φy irs (t)+n k (6)

[0062] Where Φ=diag(vec(θ))∈C Nr×Nr n k This represents the Gaussian white noise at the sensor node. The energy that the sensor node can capture during the signal reflection (SR) phase is given by the following...

[0063] E k (t)=(1-τ)ηP t L0L k MN 2 ε g (t) 2 ε f (t) 2 (7)

[0064] Based on the reflection scheme with time switching within a time slot, assume that the available energy at the beginning of time slot t is E. a (t), from which the kinetic equation for battery state update is obtained as follows:

[0065] E a (t+1)=min{E max E a (t)+e(t)-P c (t)} (8)

[0066] P c (t)=(1-τ)NP u (9)

[0067] Where e(t) is the energy absorbed in the sub-time slot τ, P c (t) is the energy consumed when the IRS reflection in this time slot is working, P u This represents the power consumption of each passive component. Furthermore, the updated remaining battery capacity is subject to capacity constraints.

[0068] E max ≥E a (t+1)≥0 (10)

[0069] Therefore, the channel capacity of the system can be expressed as:

[0070]

[0071] in This refers to the system's average noise power. Therefore, the primary objective of this invention is to achieve battery-neutral operation of the IRS and maximize the system's long-term average throughput by effectively managing the remaining battery charge through adjusting the time factor τ.

[0072] The long-term average throughput maximization model satisfies the following relationship:

[0073]

[0074] st0≤E c (n)≤E a (n+τ(n))

[0075] 0≤E a (n+1)≤E max

[0076] 0≤τ(n)≤1

[0077] E c (n)≤N r P u ;

[0078] Among them, Pu N represents the power consumption of a single reflective element in the intelligent reflective surface. r E represents the number of reflective elements in the intelligent reflective surface. c τ(n) represents the energy consumption of the intelligent reflector in time slot n, and τ(n) represents the time allocation of the energy harvesting phase of the intelligent reflector in time slot n. a (n+τ(n)) represents the remaining battery energy of the smart reflector after the energy harvesting phase in time slot n, E max E represents the maximum battery capacity of the intelligent reflective surface. a (n) represents the battery energy state starting at time slot n, Jπ represents the long-term average throughput, C(n) represents the throughput at time slot n under the execution strategy, and E a (n+1) represents the battery energy state at time slot n+1.

[0079] By adjusting the time factor τ, the remaining battery power can be effectively managed, thereby maximizing the long-term average throughput of the system.

[0080] S3. Solve the long-term average throughput maximization model based on Markov decision process, strategy iteration algorithm and objective reward function to obtain the optimal time switching strategy;

[0081] In this embodiment of the invention, step S3 includes the following sub-steps:

[0082] S31. Establish a system state model based on Markov decision process, and initialize the channel state transition matrix and battery power status.

[0083] S32. Predict the future channel state based on the system state model, and determine the energy neutrality boundary based on historical data from the sliding window.

[0084] S33. Iteratively optimize the system state model according to the strategy iteration algorithm;

[0085] S34. Determine whether the current solution is the optimal solution; if yes, output the current solution as the optimal time switching strategy; if no, return to step S33.

[0086] In this embodiment of the invention, step S31 includes the following sub-steps:

[0087] S311. Establish a state space based on the channel state of the first transmission channel, the channel state of the second transmission channel, and the remaining battery power state of the smart reflective surface.

[0088] S312. Establish an action space based on the time ratio of the battery in the absorption state and the reflection state of the intelligent reflective surface;

[0089] S313. Establish the state transition function;

[0090] S314. Establish the target reward function to obtain the system state model.

[0091] In this embodiment of the invention, the target reward function satisfies the following relationship:

[0092]

[0093] Among them, R e (t) represents the target reward function, E a (t) represents the battery energy state in time slot t, E L (t) and E U (t) represents the lower limit and upper limit of battery energy neutrality, respectively, and π represents a constant.

[0094] In this embodiment of the invention, the optimal time switching strategy satisfies the following relationship:

[0095]

[0096] Where, τ * (n) represents the optimal time switching strategy. Let R represent the average reward of the s′th state of the system state model under policy τ, γ represent the discount factor, a(t)∈A represent the action policy chosen in action space A during time slot t, and R e (t) represents the target reward function, T represents the state transition function, and a represents the action taken. S4. The intelligent reflector-assisted wireless communication system executes data transmission according to the optimal time switching strategy and updates the remaining battery power status in the intelligent reflector-assisted wireless communication system in real time.

[0097] Specifically, a system state model is established based on the Markov decision process, and an adaptive reward function R is designed. The Markov decision process is represented by a quadruple (S, A, R, T), which consists of: state space S, action space A, reward function R, and state transition function T.

[0098] The system state model consists of three parts: the channel state C from BS to IRS. g and the channel state C from IRS to sensor user Uk f And the remaining battery status E of the IRS a The state space of the entire system is S = C. g ×C f ×E a Then, within a certain time slot t, the system state is represented as: s t={C g (t),C f (t),E a (t)}, where s∈S.

[0099] The action space of the system state model mainly considers the time proportion τ of absorption and reflection within a time slot, therefore its value range is τ∈[0,1]. Action a t Represented as: a t =τ.

[0100] The system state model's state transition probability function T, to better simulate the channel gain state transition process, defines four channel gain states C = {C...} L ,C ML ,C MH ,C H} refers to low, medium-low, medium-high, and high gain, respectively. Channel state transition matrices were constructed on both sides of the IRS based on the known HCSI. and The system's state transition function T(s) t+1 |s t ,a t This describes the state s of the system at time slot t. t And execute action a t Then, it transitions to the next time slot t+1 state s. t+1 The probability distribution. Therefore, the system's state transition function T:

[0101]

[0102] For updating the remaining battery capacity, at the beginning of time slot t, the remaining usable capacity E in the battery is... a (t), and the energy consumption P of the IRS during that time slot t. c (t), the energy state E of the next time slot a (t+1) can be calculated using formula (8).

[0103] The design incorporates an objective reward function R for adaptive online energy management optimization of the IRS (Infrared Receptor System) under varying environmental conditions. The goal is to incentivize actions that promote battery energy neutrality and efficient signal transmission through proactive rewards. Specifically, when the future incident and reflected channel conditions are favorable, and the rechargeable battery maintains a low charge level, the system will employ a more aggressive strategy to transmit signal energy to the sensor nodes as much as possible. Conversely, when the future reflected channel conditions are unfavorable, and the rechargeable battery maintains a high charge level, the system will adopt a more conservative strategy to store signal energy as much as possible. This achieves online dynamic management of IRS battery energy, ensuring battery energy neutrality under channel uncertainty and maximizing the system's long-term average throughput.

[0104] In IRS battery power management, the system uses a rolling window to collect status information over a period of time, analyzes the current channel status on both sides of the IRS, and predicts channel gain. The system calculates and compares the mean and variance of the error data Γ between actual and predicted energy consumption, ensuring that the error is controlled within 95%. This ensures that even with operational errors, the IRS can still operate normally using battery power. Simultaneously, because the environmental models on both sides of the IRS change independently, the system adaptively adjusts battery limits by analyzing historical channel data, optimizing strategies while ensuring system operation. Therefore, battery power has a suitable constraint range [E] under each environmental model. L (t),E U (t)]. Therefore, a reward is given for battery power management.

[0105]

[0106] Among them, R e (t) represents the target reward function, E a (t) represents the battery energy state in time slot t, E L (t) and E U (t) represents the lower limit and upper limit of battery energy neutrality, respectively, and π represents a constant.

[0107] In Multivariate Principles (MDP), the goal is to find a policy π that maximizes the long-run expected reward. This policy π represents the average reward obtainable in the long run, starting from the system state S and acting according to policy π. It is expressed as:

[0108]

[0109] Here, R(t) is the reward obtained at time t. γ is the discount factor, representing the current value of the future reward, satisfying the boundary condition 0 < γ < 1.

[0110] Through a policy iteration algorithm, the optimal decision τ is obtained within each time block t by combining the predicted expected gain value with the remaining energy of the rechargeable and dischargeable batteries in the current IRS configuration. * value.

[0111]

[0112] Where, τ * (n) represents the optimal time switching strategy. Let R represent the average reward of the s′th state of the system state model under policy τ, γ represent the discount factor, a(t)∈A represent the action policy chosen in action space A during time slot t, and R e (t) represents the target reward function, T represents the state transition function, and a represents the action taken. S4. The intelligent reflector-assisted wireless communication system executes data transmission according to the optimal time switching strategy and updates the remaining battery power status in the intelligent reflector-assisted wireless communication system in real time.

[0113] In this embodiment of the invention, the intelligent reflector-assisted wireless communication system makes a decision τ* before the arrival of time slot t. After the actual time slot t channel state arrives, the intelligent reflector-assisted wireless communication system will execute data transmission according to the optimal time switching strategy and update the remaining battery power status in real time to ensure the system operates with neutral battery power and improve system performance.

[0114] Compared to existing technologies, this invention utilizes a sliding window to update historical data in real time within each time slot. This historical data includes the incident and reflected channel states of past time slots, the difference between energy harvesting and reflected energy consumption of the intelligent reflector, and the battery energy status. Using the energy difference data and battery energy status, the system obtains an initial energy boundary. An energy boundary adjustment factor is calculated based on the channel states of past time slots, and this factor is used to adjust the initial boundary to obtain an energy-neutral target boundary. An energy optimization method based on Markov processes and online time-switching strategies is employed. According to the designed target reward function, when the overall channel conditions on both sides of the intelligent reflector are poor, the system stores energy through an online time-switching strategy. When the overall channel conditions of the intelligent reflector are good, efficient data transmission is implemented. An optimal time-switching strategy based on the energy-neutral boundary is obtained through strategy iteration algorithm optimization. Upon the arrival of the actual channel environment, the system executes its decisions and updates the remaining battery power status. This allows the intelligent reflector to adaptively maintain a certain amount of remaining power under different dynamic environments, significantly enhancing the reliability and efficiency of the communication system, while ensuring the continuity and stability of communication in dynamically changing wireless network environments. This enables the battery energy neutrality of the intelligent reflective surface and maximizes the long-term average throughput of the system, thereby greatly improving the overall performance of the wireless communication system.

[0115] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0116] The embodiments of the present invention have been described above with reference to the accompanying drawings. The disclosed embodiments are merely preferred embodiments of the present invention. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many equivalent changes in form without departing from the spirit and scope of the claims of the present invention, and all such changes are within the protection scope of the present invention.

Claims

1. A method for online throughput optimization of a wireless communication system based on a smart reflector, characterized in that, Includes the following steps: S1. Construct a battery-powered intelligent reflective surface-assisted wireless communication system; S2. Based on the online time switching strategy, the intelligent reflector-assisted wireless communication system is divided into an energy harvesting stage and a signal reflection stage, and a long-term average throughput maximization model is established with the goal of maximizing the long-term average throughput of the intelligent reflector-assisted wireless communication system and achieving battery energy neutrality. S3. Solve the long-term average throughput maximization model based on Markov decision process, strategy iteration algorithm and objective reward function to obtain the optimal time switching strategy; S4. The intelligent reflective surface-assisted wireless communication system performs data transmission according to the optimal time switching strategy and updates the remaining power status of the battery in the intelligent reflective surface-assisted wireless communication system in real time. Step S3 includes the following sub-steps: S31. Establish a system state model based on Markov decision process, and initialize the channel state transition matrix and battery power status. S32. Predict the future channel state based on the system state model, and determine the battery energy neutrality boundary based on historical data from the sliding window. S33. Iteratively optimize the system state model according to the strategy iteration algorithm; S34. Determine whether the current solution is the optimal solution; if yes, output the current solution as the optimal time switching strategy; if no, return to step S33.

2. The online throughput optimization method for a wireless communication system based on a smart reflector as described in claim 1, characterized in that, The intelligent reflector-assisted wireless communication system includes a base station, an intelligent reflector equipped with a rechargeable battery, and multiple single-antenna sensor nodes.

3. The online throughput optimization method for a wireless communication system based on a smart reflector as described in claim 1, characterized in that, The long-term average throughput maximization model satisfies the following relationship: ; in, This represents the power consumption of a single reflective element in the intelligent reflective surface. N r This indicates the number of reflective elements in the intelligent reflective surface. The energy consumption of the intelligent reflective surface is represented by the value in time slot n. This represents the time allocation for the energy harvesting phase of the intelligent reflector in time slot n. This represents the remaining battery energy of the smart reflector after the energy harvesting phase in time slot n. This indicates the maximum battery capacity of the intelligent reflective surface. , This represents the long-term average throughput. , .

4. The online throughput optimization method for a wireless communication system based on a smart reflector as described in claim 1, characterized in that, Step S31 includes the following sub-steps: S311. Establish a state space based on the channel state of the first transmission channel, the channel state of the second transmission channel, and the remaining battery power state of the smart reflective surface. S312. Establish an action space based on the time ratio of the battery in the absorption state and the reflection state of the intelligent reflective surface; S313. Establish the state transition function; S314. Establish the target reward function to obtain the system state model.

5. The online throughput optimization method for a wireless communication system based on a smart reflector as described in claim 4, characterized in that, The target reward function satisfies the following relationship: in, Describes the target reward function. This represents the battery energy state at time slot t. and These are the lower limit and the upper limit of battery energy neutrality, respectively. Represents a constant.

6. The online throughput optimization method for a wireless communication system based on a smart reflector as described in claim 5, characterized in that, The optimal time switching strategy satisfies the following relationship: in, This represents the optimal time switching strategy. The first part of the system state model represents the... The average reward of a state under policy τ Indicates the discount factor. , , 。

Citation Information

Patent Citations

  • Resource allocation method of wireless sensor network assisted by intelligent reflecting surface

    CN113709687A

  • Energy management method for energy collection wireless sensor

    CN117412367A