Short packet transmission method and system based on irregular IRS assisted backscatter communication

By constructing a downlink short packet communication scenario assisted by an irregular active intelligent metasurface, and using a deep reinforcement learning model to optimize the transmission beam of the base station and the intelligent metasurface, the problem of insufficient short packet transmission capability of 6G IoT devices is solved, achieving efficient short packet transmission and system capacity improvement.

CN116980931BActive Publication Date: 2026-07-24NANJING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
NANJING UNIV OF POSTS & TELECOMM
Filing Date
2023-05-12
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

In 6G scenarios, IoT devices have insufficient short packet transmission capabilities, especially under limited conditions. How can we improve the transmission efficiency and reliability of IoT devices?

Method used

We construct a downlink short packet communication scenario assisted by an irregular active intelligent metasurface. We optimize backscatter communication through a deep reinforcement learning model that determines policy gradients, and jointly optimize the transmission beam of the base station and the intelligent metasurface to maximize the short packet transmission rate.

Benefits of technology

It improves the short packet transmission rate and system capacity of IoT devices, enhances information transmission from base stations to users, and meets the requirements for ultra-reliable and low-latency communication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116980931B_ABST
    Figure CN116980931B_ABST
Patent Text Reader

Abstract

The application discloses a kind of short packet transmission methods and systems based on irregular IRS assisted backscattering communication, method includes: constructing irregular active intelligent metasurface assisted downlink short packet communication scene;Under the power constraint of base station transmitting power, backscattering signal-to-interference-and-noise ratio, irregular active intelligent metasurface, construct the optimization problem model of maximum backscattering short packet transmission rate;The optimization problem model is jointly optimized using the deep reinforcement learning model based on deep deterministic policy gradient, an optimization algorithm is proposed, the maximum neural network feedback under the optimal state is obtained, and then the highest transmission rate is obtained;The present application considers that irregular active IRS can further develop the spatial degree of freedom of reflecting unit, and then improve the system capacity, irregular active IRS backscattering communication sends information to user, enhances the information transmission from BS to user, increases the system transmission rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of wireless communication technology, and in particular to a short packet transmission method and system based on irregular IRS-assisted backscatter communication. Background Technology

[0002] To fully realize a smart information society, large-scale device connectivity, data-driven services, and autonomous systems, the Internet of Things (IoT) in 6G scenarios will face more stringent requirements. For example, the Ultra-Reliable Low-Latency Communications (URLLC) standard will be more stringent. Although URLLC has already been introduced and used in 5G-based IoT applications, it still needs improvement to support emerging large-scale IoT applications in 6G scenarios. In the future, URLLC will be used in smart factories to automate critical processes, such as automated smart manufacturing and remote robot control. Compared to traditional wired connections, using URLLC technology can optimize operating costs with extremely low latency and ultra-high reliability, achieving a reliability of up to 99.9999% and less than 10... -5 The block error rate. Furthermore, the coverage of IoT networks in 6G scenarios will be further expanded compared to current transmission scenarios, encompassing more network types and exhibiting greater heterogeneity.

[0003] Key technologies for IoT in 6G scenarios include: edge intelligence, intelligent reconfigurable surfaces (IRS), integrated space-ground communication, terahertz communication, massive URLLC connections, and blockchain technology. IRS primarily consists of passive reflective elements with artificial planar structures, each component reflecting incident electromagnetic waves in a software-defined manner using electronic circuitry. IRS supports wireless system design optimization by facilitating signal propagation, channel modeling, and data acquisition, thereby making the intelligent radio environment favorable for IoT applications. IRS can simultaneously enhance the signals collected by serving base stations (BS) in multi-cell heterogeneous IoT networks, reducing cell interference between numerous IoT devices. IRS is a promising solution for the future development of the IoT industry in 6G scenarios. With the support of key technologies, 6G is expected to facilitate the implementation of various emerging IoT applications, including healthcare IoT, autonomous driving in vehicles, the drone industry, satellite IoT, and industrial IoT. For example, in the industrial IoT field, the limited resources and random deployment of IoT sensors in future industrial scenarios may lead to problems such as poor information transmission quality, low response efficiency, and a low number of accessible devices. IRS can provide high-quality information transmission services for industrial IoT in 6G scenarios.

[0004] However, the implementation of IoT applications in 6G scenarios still faces many challenges. As is well known, wireless network resources have always been scarce, and how to further improve the short packet transmission capability of IoT devices under limited conditions has been a hot research direction. Among these, IRS (Indirect Redirect Radio) represents the future of IoT communication in 6G networks. Recently, a symbiotic radio system to support passive IoT has been proposed, in which a backscattering device (BD), also known as an IoT device, coexists with the main transmission. The main transmitter assists in both the main transmission and BD transmission, while the main receiver decodes information from both the main transmitter and BD. The symbol period of the BD transmission is assumed to be equal to or greater than the symbol period of the main transmission, thus forming a symbiotic situation that allows the BD to transmit data and improves spectrum utilization. Many studies have also combined IRS and backscattering communication, proposing a novel IRS-assisted multiple-input multiple-output (MIMO) radio symbiotic system. In this system, the IRS acts as a secondary transmitter, sending information to a multi-antenna secondary receiver through cognitive backscattering communication, while intelligently reconfiguring the wireless environment, enhancing the primary transmission from the multi-antenna main transmitter to the multi-antenna main receiver. In cellular networks, IRS can also be used to assist cellular link communication while transmitting additional data for various IoT applications, which aligns with the concept of backscattering. Existing technology studies an uplink massive MIMO backscatter radio system based on IRS, where each IRS acts as an IoT device, enhancing the primary transmission from nearby users to the BS, while simultaneously transmitting its own information to the BS via backscatter modulation. By embedding environmental sensors on the IRS, the system can transmit locally collected environmental data to the BS via IoT and assist in primary transmission. Another approach is a broadcast-based IRS-licensed backscatter radio, where the BS broadcasts signals to multiple master receivers with the help of the IRS, and the IRS also transmits information to IoT receivers via the broadcast signals. While these studies offer varying degrees of reference value, they do not consider how to ensure high-reliability, low-latency communication in the context of backscattering. Furthermore, the current irregularity of IRS is a worthwhile research direction, and further improvements in transmission efficiency through additional spatial degrees of freedom could be explored. Summary of the Invention

[0005] The purpose of this section is to outline some aspects of embodiments of the present invention and to briefly describe some preferred embodiments. Simplifications or omissions may be made in this section, as well as in the abstract and title of this application, to avoid obscuring the purpose of these documents; however, such simplifications or omissions should not be construed as limiting the scope of the invention.

[0006] In view of the above-mentioned problems, the present invention is proposed.

[0007] Therefore, the technical problem solved by the present invention is to further improve the short packet transmission capability of Internet of Things devices under limited conditions.

[0008] To solve the above-mentioned technical problems, the present invention provides the following technical solution:

[0009] In a first aspect, embodiments of the present invention provide a short packet transmission method based on irregular IRS-assisted backscatter communication, comprising:

[0010] Constructing a downlink short packet communication scenario assisted by irregular active intelligent metasurfaces;

[0011] Under the constraints of base station transmit power, backscatter signal-to-interference-plus-noise ratio, and power of irregular active intelligent metasurface, an optimization problem model is constructed to maximize the transmission rate of backscatter short packets.

[0012] The optimization problem model is jointly optimized using a deep reinforcement learning model based on a deep deterministic policy gradient. An optimization algorithm is proposed to obtain the maximum neural network feedback in the optimal state, thereby achieving the highest transmission rate.

[0013] As a preferred option for short packet transmission methods based on irregular IRS-assisted backscatter communication, wherein:

[0014] The downlink short packet communication scenario assisted by the construction of irregular active intelligent metasurfaces includes:

[0015] The communication scenario includes a base station, an irregular active intelligent metasurface, and a short-packet communication user. The base station has a multi-antenna structure, and the receiver has a single antenna. The system transmits information to the short-packet user via backscatter communication through the irregular active intelligent metasurface. The transmitting base station is equipped with multiple antennas (M>1), and the low-latency users are all single-antenna users. The intelligent metasurface consists of… It consists of active reflective units, in which N reflective units are distributed in N s N in each grid space s >N, where the grid constraint with a fixed grid spacing is 1, and the grid spacing between adjacent grid points is set to half the wavelength of the carrier frequency, wherein the irregular active intelligent metasurface can be realized by selecting a certain number of reflective units from all grid points; the intelligent metasurface assists the transmission of the main signal from the base station to the user, and at the same time modulates its own signal onto the incident signal through backscattering technology.

[0016] As a preferred option for short packet transmission methods based on irregular IRS-assisted backscatter communication, wherein:

[0017] The downlink short packet communication scenario assisted by the construction of irregular active intelligent metasurfaces also includes:

[0018] This represents the channel from the base station to the smart metasurface. This represents the channel from the base station to the receiving user. This represents the channel from the smart metasurface to the receiving user; each channel consists of two parts: a large-scale fading component and a small-scale fading component. The large-scale fading is distance-dependent and is modeled as follows:

[0019]

[0020] Where d represents the distance from the signal transmission end to the receiving end, β represents the path loss per 1m of the reference distance, and γ e d represents the path loss coefficient. br d represents the distance from the base station to the receiver. g d represents the distance from the base station to the smart metasurface. ir This represents the distance from the smart metasurface to the receiver; the small-scale fading of each channel follows a Ricean fading channel model, which consists of a line-of-sight (LoS) component and a non-line-of-sight (NLoS) component:

[0021]

[0022] Where k br Represents Rice factor, and Representing channel h respectively d The visual range and the non-visual range; It follows a complex Gaussian distribution with zero mean and unit variance, while the line-of-sight component... The vector representation model is shown below:

[0023]

[0024] in d a This represents the space of the antenna, where λ is the wavelength and θ is the angle between the two wavelengths. AoA And θAoD represents the angle of arrival of the receiver and the angle parameter of the transmission angle of the base station.

[0025] As a preferred option for short packet transmission methods based on irregular IRS-assisted backscatter communication, wherein:

[0026] The optimization problem model for maximizing the backscatter short packet transmission rate includes:

[0027] The optimization problem, which aims to jointly optimize the transmission beam at the base station and the reflected beam at the smart metasurface to maximize the short packet transmission rate, is constrained by the transmission power at the base station, the minimum signal-to-noise ratio of backscatter transmission, and the active irregular smart metasurface constraint. The problem can be expressed as follows:

[0028] P:

[0029] stC1:|w| 2 ≤P T

[0030] C2:γ c ≥SINR c

[0031] C3:

[0032] C4:

[0033] C5:1 T Z = N

[0034] Where w is the transmission beam vector at the base station, and P in constraint C1 T Indicates the base station's transmit power, γ c The average signal-to-noise ratio is given by ρ, where ρ is the amplification factor constraint, and z is the average i =1 indicates that a smart metasurface reflective unit is deployed at the i-th grid point, z i =0 indicates that no smart metasurface reflection unit is deployed at the i-th grid point. Let Z = diag(z), and G represent the channel from the base station to the smart metasurface. When the smart metasurface transmits symbol 1, it is represented as Θ, and when it transmits symbol -1, it is represented as -Θ. C3 represents the noise enhancement power; C4 and C5 represent the maximum enhancement power constraint at the smart metasurface; C4 and C5 represent the constraints on the reflective units of the irregular smart metasurface panel; there are a total of N for the irregular smart metasurface portion. s The reflective unit is used, but the actual number of units used is only N.

[0035] As a preferred option for short packet transmission methods based on irregular IRS-assisted backscatter communication, wherein:

[0036] The joint optimization includes:

[0037] The input channels G from the base station to the smart metasurface, and h from the smart metasurface to the URLLC user. r direct channel h from base station to URLLC user d Output the optimal action space a = [ρ, w, Z, Θ], and the real-time reward r.

[0038] As a preferred option for short packet transmission methods based on irregular IRS-assisted backscatter communication, wherein:

[0039] The joint optimization also includes:

[0040] First, initialization is performed; specifically, the online actor network μ(θ) μ |s), online critic network Q(θ) Q |s,a), target critic network Q′(θ) Q′ |s,a), target actor network μ′(θ) μ′ |s) and satisfy θ μ′ =θ μ θ Q′ =θ Q Experience replay pool Small batch Online network learning rate ζ Q and ζ μ ;

[0041] Begin the outer loop, which includes performing the following operations for each i = 1, 2, ..., I: obtain G (i) , as well as Initialize the amplification factor ρ (0) Set w to 1 and initialize it with a unit vector. (0) Initialize Z with the identity matrix (0) ,Θ (0) Initialization was handled using the OU noise process. Explore actions and obtain the initial state s (1) .

[0042] As a preferred option for short packet transmission methods based on irregular IRS-assisted backscatter communication, wherein:

[0043] The joint optimization also includes:

[0044] Start the inner loop: For each time step t = 1, 2, ..., T, perform the following operations: based on a (t) Calculate [P] T (t) ,P I (t) [Get instant reward feedback r] (t) And the new state s of the next time slot. (t+1) The experience gained (s) (i) ,a (i) ,r (i) ,s (i+1) Stored in the experience replay pool In the middle, from the experience replay pool Medium sampling A random mini-batch of transitions is used to set the target Q-network using the following formula:

[0045]

[0046] in γ∈(0,1] represents the discount factor, and the online critic network Q(θ) Q The loss function L(θ) of |s,a) Q The presentation format is as follows:

[0047]

[0048] Policy gradient for building an online actor network Represented as:

[0049]

[0050] Update the online critic and the online actor network separately, as shown below:

[0051]

[0052]

[0053] Where ζ Q and ζ μ This represents the network learning rate, used to update the parameters of the online critic network and the online actor network, respectively.

[0054] The target critic network and the target actor network are soft-updated separately, as follows:

[0055] θ Q′ ←τ Q θ Q -(1-τ Q )θ Q′

[0056] θ μ′ ←τ μ θ μ -(1-τ μ )θ μ′

[0057] Where τ Q and τ μ It also represents the network learning rate, which is used to update the target critic network and the target actor network, respectively;

[0058] Let s (t) =s (t1) The inner loop ends, and the outer loop ends.

[0059] Secondly, embodiments of the present invention provide a short packet transmission system based on irregular IRS-assisted backscatter communication, characterized in that it includes:

[0060] The communication scenario building module is used to construct downlink short packet communication scenarios assisted by irregular active intelligent metasurfaces;

[0061] The optimization problem modeling module is used to construct an optimization problem model that maximizes the backscatter short packet transmission rate under the power constraints of base station transmit power, backscatter signal-to-interference-plus-noise ratio, and irregular active intelligent metasurface.

[0062] The joint optimization module is used to jointly optimize the optimization problem model using a deep reinforcement learning model based on deep deterministic policy gradient, propose an optimization algorithm, obtain the maximum neural network feedback in the optimal state, and thus obtain the highest transmission rate.

[0063] Thirdly, embodiments of the present invention provide a computing device, including:

[0064] Memory and processor;

[0065] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the one or more programs are executed by the one or more processors, the one or more processors implement the short packet transmission method based on irregular IRS-assisted backscatter communication as described in any embodiment of the present invention.

[0066] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the aforementioned short packet transmission method based on irregular IRS-assisted backscatter communication.

[0067] The beneficial effects of this invention are as follows: This invention further considers the distribution of IRS reflective units and studies the problem of ultra-reliable low-delay backscattering transmission communication based on irregular active IRS assistance. Considering that irregular active IRS can further develop the spatial degrees of freedom of reflective units, thereby improving system capacity, irregular active IRS backscattering communication sends information to users, enhances information transmission from the BS to the user, and increases the system transmission rate. Attached Figure Description

[0068] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:

[0069] Figure 1 This is an overall flowchart of the short packet transmission method based on irregular IRS-assisted backscatter communication described in the first embodiment of the present invention;

[0070] Figure 2 This is a system model diagram of the short packet transmission method based on irregular IRS-assisted backscatter communication described in the first embodiment of the present invention.

[0071] Figure 3 This is a DRL framework diagram of the short packet transmission method based on irregular IRS-assisted backscatter communication described in the first embodiment of the present invention;

[0072] Figure 4 This is an iterative convergence graph from a simulation example of the short packet transmission method based on irregular IRS-assisted backscatter communication described in the second embodiment of the present invention.

[0073] Figure 5 This is a schematic diagram illustrating the transmission behavior of the degree of irregular IRS with respect to the decoding error probability in a simulation example of the short packet transmission method based on irregular IRS-assisted backscatter communication according to the second embodiment of the present invention.

[0074] Figure 6 This is a schematic diagram illustrating the relationship between the degree of irregularity of IRS and the horizontal distance in a simulation example of the short packet transmission method based on irregular IRS-assisted backscatter communication described in the second embodiment of the present invention. Detailed Implementation

[0075] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.

[0076] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0077] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.

[0078] This invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of this invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not adhering to the usual scale. Furthermore, the schematic diagrams are merely examples and should not be construed as limiting the scope of protection of this invention. In actual fabrication, the three-dimensional spatial dimensions of length, width, and depth should be included.

[0079] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.

[0080] Unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" in this invention should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; similarly, they can refer to mechanical connections, electrical connections, or direct connections, or indirect connections through an intermediate medium, or internal connections between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0081] Example 1

[0082] Reference Figures 1-3 This is the first embodiment of the present invention, which provides a short packet transmission method based on irregular IRS-assisted backscatter communication, comprising:

[0083] S1: Construct a downlink short packet communication scenario assisted by an irregular active IRS. This scenario includes a base station, an irregular active IRS, and a short packet communication user.

[0084] Furthermore, in the downlink network transmission scenario assisted by irregular active IRS, this scenario mainly includes a power-transmitting BS, an irregular active IRS, and URLLC users. The transmitting BS is equipped with multiple antennas (M>1), while low-latency users are all single-antenna users. The IRS is... It consists of active reflective units, in which N reflective units are distributed in N s In N grid spaces sTo simplify the description of the proposed concept, this invention assumes a grid constraint of 1 with a fixed grid spacing. Without loss of generality, the grid spacing between adjacent grid points is assumed to be half the wavelength of the carrier frequency. The irregular active IRS can be implemented by selecting a certain number of reflective elements from all grid points. The IRS assists the main signal transmitted from the BS to the user, while simultaneously modulating its own signal onto the incident signal using backscattering technology.

[0085] S2: In the communication scenario, the base station has a multi-antenna structure and the receiver has a single antenna. Considering that irregular active IRS can further improve the spatial degree of freedom of the reflector unit and thus improve the system capacity, the system sends information to short packet users through irregular active IRS backscatter communication, which enhances the transmission signal from the BS to the user.

[0086] Furthermore, consider a block-flat fading channel model, where the channel coefficients remain constant within a block period but may change as you move from one block period to another. For example... Figure 2 As shown, This indicates the channel from the BS to the IRS. This indicates the channel from the BS to the receiving user. This represents the channel from the IRS to the receiving user. Each channel consists of two parts: a large-scale fading component and a small-scale fading component. The large-scale fading, which is distance-dependent, can be modeled as follows:

[0087]

[0088] Where d represents the distance from the signal transmission end to the receiving end, β represents the path loss per 1m of the reference distance, and γ e Let d represent the path loss coefficient. br d represents the distance from the BS to the receiver. g d represents the distance from BS to IRS. ir This represents the distance from the IRS to the receiver. Without loss of generality, the small-scale fading of each channel follows a Ricean fading channel model, where h is taken as a direct-connected channel. d Taking one example, the other two channels are analyzed in the same way. The Ricean fading channel model consists of a line-of-sight (LoS) component and a non-line-of-sight (NLoS) component:

[0089]

[0090] Where k br Represents Rice factor, and Representing channel h respectively d The visual distance and the non-visual distance portion. It follows a complex Gaussian distribution with zero mean and unit variance, while the line-of-sight component... The vector representation model is shown below.

[0091]

[0092] in d a This represents the space of the antenna, where λ is the wavelength and θ is the angle between the two wavelengths. AoA and θ AoD This represents the angle of arrival at the receiver and the angle of transmission of the BS.

[0093] S3: Under the constraints of BS transmit power, backscatter signal-to-interference-plus-noise ratio, and power of irregular active IRS, construct an optimization problem model to maximize the transmission rate of backscatter short packets.

[0094] Furthermore, the transmission symbol at BS is denoted as x(l), where Let be the transmitted beam vector at BS, then the transmitted signal at BS can be written as wx(l). In an active IRS, each IRS unit has a reflection enhancement function, so the reflection unit coefficient ρ i >1, for ease of processing, we assume here. so It needs to be multiplied by an amplification factor ρ, therefore Similarly, a magnification factor ρ is needed, Θ = ρΦ. For active IRS, a non-negligible issue also needs to be considered: each IRS has enhanced noise. It follows an independent circularly symmetric complex Gaussian distribution. in This represents the noise enhancement power. To ensure continuous phase change, a varactor diode is loaded for each reflective element. Backscattering is used, and a binary phase-shift keying modulation scheme (c∈{-1,1}) is employed during the two-stage transmission. Therefore, the backscattered signal at the IRS is cΘ, which yields the k-th IRS reflective unit. That is, when the IRS sends symbol 1, it is represented as Θ, and when it sends symbol -1, it is represented as -Θ.

[0095] The backscatter communication method uses backscatter symbols, where the backscatter symbol c consists of L periods of active transmission symbols, i.e., T c =LT x Let x(1), x(2), ..., x(L) represent active transmission symbols contained in a symbol c. There are a total of N for the irregular IRS portion. S There are N reflective units, but the actual number of units used is only N, let Where z i∈{1,0} indicates whether the IRS cell is deployed at the i-th grid point, z i =1 indicates that an IRS reflector unit is deployed at the i-th grid point, z i =0 indicates that no IRS reflection unit is deployed at the i-th grid point. Here, let Z = diag(z). The addition of irregular IRS enhances the transmission capacity by increasing the spatial degrees of freedom of the IRS. In summary, the reflected signal at the IRS can be expressed as ρZΘGwx(L)c, so the signal at the receiving end is expressed as:

[0096]

[0097] The signal receiver uses Maximum Ratio Combining (MRC) to decode the x(l) signal, so the average SINR during x(l) decoding is as follows:

[0098]

[0099] Assuming x(l) can be correctly decoded, the directly connected signal is subtracted at the receiving end, and MRC is applied again to the remaining signal. Then the SINR of the decoded c is...

[0100]

[0101] Therefore, when decoding c, the average SINR is:

[0102]

[0103] Therefore, we can obtain the irregular IRS backscatter short packet transmission R. x As shown below:

[0104]

[0105] Where a = log₂e, ε < 10 -5 Q -1 (·) denotes the inverse of the Gaussian Q-function, and we have This invention aims to jointly optimize the transmission beam at the BS and the reflection beam at the IRS, thereby maximizing the short packet transmission rate to meet user needs. However, due to constraints such as the transmission power at the BS, the minimum SINR of backscatter transmission, and the active irregular IRS, the corresponding optimization problem can be written in the form of the following problem (9):

[0106]

[0107] Where P in constraint C1 TThe signal represents the transmit power of the BS. Constraint C3 represents the maximum boost power constraint at the IRS, which is actually much smaller than that of a conventional RF amplifier due to limited amplification power. C4 and C5 represent the constraints of the irregular IRS panel reflector unit.

[0108] S4: Considering the non-convexity of the optimization problem, a deep reinforcement learning model based on DDPG is used to jointly optimize the problem in S3. An optimization algorithm is proposed to obtain the maximum neural network feedback in the optimal state and thus obtain the highest transmission rate.

[0109] Furthermore, this step utilizes the Deep Reinforcement Learning (DRL) framework:

[0110] Deep Reinforcement Learning (DRL) is a fusion of deep neural networks and reinforcement learning. It is primarily used to solve Markov Decision Processes (MDRs), which are characterized by large-scale action and state spaces that are difficult to solve effectively using traditional reinforcement learning methods. Several key elements need to be considered in deep reinforcement learning:

[0111] 1) State: State space, This represents a series of values ​​observed in the environment at time slot t, where It represents the set of state spaces in the environment, where the environment can be regarded as a system model.

[0112] 2) Action: Action space This represents the state value s based on the observation at time slot t. (t) Below is the set of actions chosen according to strategy π, where Represents the action space.

[0113] 3) Policy: Policy space, π(a (t) |s (t) ) represents based on environmental state s (t) The action decision made by the agent in time slot t is a (t) .

[0114] 4) Reward: Reward function This is a single-step feedback value, and the prerequisite for this feedback value is that the agent is in state s. (t) Execute a (t) The instantaneous feedback value in time slot t represents the state s. (t) Execute a (t) The performance of the agent is such that it can adjust its feedback r accordingly through this instantaneous feedback.

[0115] 5) State-Action Value Function: The state-action value function Q is defined as follows: This value indicates the execution of action a. (t) The subsequent evaluation of cumulative rewards, where θ represents the parameters of the neural network based on policy π, which is the key information that the neural network needs to train. Additionally... This represents the cumulative reward value, and γ∈(0,1] represents the discount factor.

[0116] It should be noted that DDPG is highly effective in solving problems in continuous action spaces due to its fusion of the advantages of actor-critic networks, policy gradient determination, and deep Q-networks. The main logic is as follows: First, the policy function μ and Q-function are constructed using an AC network; then, DDPG uses two neural networks as simulations of the policy function μ and Q-function, respectively. One network is a policy-based actor neural network, and the other is a value-based critic network. Regarding network optimization, to address the instability during network training—specifically, the estimated Q-value tends to diverge—the target network is used in the deep Q-network. Therefore, it includes four neural networks: an online actor network μ(θ) and a deep Q-network. μ |s), target actor network μ′(θ) μ′ |s), online critic network Q(θ) Q |s,a) and the target critic network Q′(θ) Q′ |s,a), and θ i / i′ (i = μ / Q) represents the network parameters. The target network and the online network have the same network structure but different parameters, and there is also an experience replay pool. It is a tuple consisting of the state space, action space, reward function, and empirical parameters such as the next time slot state, represented as (s (t) ,a (t) ,r (t) ,s (t+1) This tuple was originally used in deep Q-networks. Here, its use in DDPG networks can interfere with the correlation between experiences, thereby making the training of neural networks more effective.

[0117] Since reinforcement learning obtains the optimal selection strategy through continuous exploration, in order to improve the efficiency of exploration, the exploration strategy needs to be optimized. Add a noise processing procedure Random sampling small batch (s) (i) ,a (i) ,r (i) ,s (i+1) ),and They are all from the experience replay pool Since the samples are randomly selected, the target Q value can be expressed in the following form:

[0118]

[0119] in γ∈(0,1] represents the discount factor, and the online critic network Q(θ) Q The loss function of |s,a) is expressed as follows:

[0120]

[0121] Known as the performance objective function, it is used in deep policy gradient networks to measure the specific policy μ. θ (s), where υ μ (s) represents the policy μ and action function a = μ θ The state distribution function of (s). When using offline policy training, the policy gradient is obtained through this representation. Based on the Monte Carlo method, it is possible to obtain [data] by using small batches of data. The unbiased estimate. Therefore, the policy gradient of the online actor network can be rewritten in the following form:

[0122]

[0123] The online critic network and the online actor network update their network parameters according to the following formula:

[0124]

[0125]

[0126] Where ζ Q and ζ μ This represents the network learning rate, used to update the parameters of the online critic network and the online actor network, respectively.

[0127] The target network then updates its parameters via soft updates, and can be rewritten as follows:

[0128] θ Q′ ←τ Q θ Q -(1-τ Q )θ Q′ (15)

[0129] θ μ′ ←τ μ θ μ -(1-τ μ)θ μ′ (16)

[0130] Where τ Q and τ μ It also represents the network learning rate, which is used to update the target critic network and the target actor network, respectively.

[0131] It should also be noted that in the DDPG-based joint optimization algorithm, the amplification factor is constant, the BS transmit beam, irregular reflection beam, and irregular IRS distribution are large-scale continuous matrix vectors. The entire communication scenario is treated as the DRL interaction environment, and the controllers controlling the reflection factor, irregular vector, and active / passive beams are treated as intelligent agents, specifically as follows: Figure 3 As shown.

[0132] 1) State: Represented as

[0133] s (t) =[P T (t-1) ,P I (t-1) ,G,h r ,h d ,a (t-1) (17)

[0134] Where P T (t-1) This represents the maximum transmit power of the BS in time slot t-1, P I (t-1) This represents the maximum enhancement power constraint for the irregular IRS in time slot t-1, where G represents the channel matrix from the BS to the irregular IRS, and h represents... r h represents the reflection channel matrix from an irregular IRS to a URLLC user. d This represents the direct channel matrix from the BS to the URLLC user, a (t-1) This indicates the action selection in time slot t-1.

[0135] 2) Action: Represented as

[0136] a (t) =[ρ (t) ,w (t) Z (t) ,Θ (t) (18)

[0137] a (t) The parameters within represent the reflection coefficient, active beam, irregular IRS arrangement, and passive beam of time slot t, respectively, which are the optimization terms in the problem.

[0138] 3) Reward: The reward function represents the optimization objective. The objective of this step is to maximize the user rate for short packets, so it is expressed in the following form:

[0139]

[0140] And there are:

[0141]

[0142]

[0143]

[0144]

[0145] In formula (19), part 0 represents the optimized target low-latency user transmission rate, part 1 represents the minimum SINR requirement for backscattered symbols in the irregular IRS at time slot t, part 2 represents the active boost power requirement for the irregular IRS at time slot t, part 3 represents whether reflective elements are placed in the grid points of the irregular IRS, and part 4 represents the required number of reflective elements that can be placed in all grid points of the irregular IRS. It is a normal number obtained through continuous training, used to measure utility and loss. The addition of formulas (20) and (21) applies backscattering and active IRS consumption requirements to low-latency user rates, respectively. The addition of formulas (22) and (23) ensures the irregular IRS setting with respect to grid points. If these four requirements are met in time slot t, then... This indicates that the constraints of the optimization problem are guaranteed and the reward function is not penalized.

[0146] In the DRL framework's Agent, all actor neural networks and critic neural networks include one input layer, one output layer, and two hidden layers. Their network structure is fully connected. (t) and s (t) These represent the input and output of the actor network, respectively. In the critic network, a... (t) and s (t) The input is Q, and the output is Q. This invention uses batch regularization to address the issue of varying data distribution at each layer, making the new data distribution more suitable for the true data distribution. Batch regularization also speeds up model training by ignoring dropout. Using L1 and L2 regularization improves training accuracy, and the transmit power constraint is handled separately at the output processing layer.

[0147] As described in S4, the entire DRL framework involves a Markov decision process, which, in each episode's time slot t, is based on a (t-1) =[ρ (t-1) ,w (t-1) Z (t-1) ,Θ (t-1) ], a (t-1) The calculated [P] T (t-1) ,P I (t-1) ] and channel information [G,h] obtained from the entire communication environment r ,h d An intelligent agent can then construct state s (t) Then with s (t) As an incentive based on policy θ μ The agent should provide a corresponding action a. (t) =[ρ (t) ,w (t) Z (t) ,Θ (t) Next, through a (t) And the reward r issued by the environment (t) That is, optimizing the objective function to obtain the low-latency user transmission rate [P] T (t) ,P I (t) The new state s of the next time slot (t+1) =[P T (t) ,P I (t) ,G,h r ,h d ,a (t) This process is repeated iteratively. The detailed algorithm flow based on DDPG is shown below:

[0148] Input: Channel G from BS to IRS, channel h from IRS to URLLC user r Direct channel h from BS to URLLC user d ;

[0149] Output: Optimal action space a = [ρ, w, Z, Θ], real-time reward r;

[0150] 1: Initialization: Online actor network μ(θ) μ |s), online critic network Q(θ) Q |s,a), target critic network Q′(θ) Q′ |s,a), target actor network μ′(θ) μ′ |s) and satisfy θμ′ =θ μ θ Q′ =θ Q Experience replay pool Small batch Online network learning rate ζ Q and ζ μ ;

[0151] 2: Start the outer loop: for episode i = 1, 2, ..., I do;

[0152] 3: Obtain as well as

[0153] 4: Initialize the amplification factor ρ (0) Set w to 1 and initialize it with a unit vector. (0) Initialize Z with the identity matrix (0) ,Θ (0) ;

[0154] 5: Use OU noise process to handle initialization Explore actions;

[0155] 6: Obtain the initial state s (1) ;

[0156] 7: Start the inner loop: for time step t = 1, 2, ..., T do;

[0157] 8: Based on a (t) Calculate [P] T (t) ,P I (t) [Get instant reward feedback r] (t) And the new state s of the next time slot. (t+1) ;

[0158] 9: The experience gained (s) (i) ,a (i) ,r (i) ,s (i+1) Stored in the experience replay pool middle;

[0159] 10: From the experience replay pool Medium sampling A transitional random mini-batch;

[0160] 11: Set the target Q-network using formula (10);

[0161] 12: Construct the loss function L(θ) of the online critic network using formula (11). Q );

[0162] 13: Construct the policy gradient of the online actor network using formula (12)

[0163] 14: Update the online critic and online actor networks using formulas (13) and (14) respectively;

[0164] 15: Use formulas (15) and (16) respectively to soft update the target critic network and the target actor network;

[0165] 16: Let s (t) =s (t+1) ;

[0166] 17: Inner loop ends;

[0167] 18: The outer loop ends.

[0168] It should also be noted that this invention further considers the distribution of IRS reflective units and studies the problem of ultra-reliable low-delay backscatter transmission communication based on irregular active IRS assistance. Considering that irregular active IRS can further develop the spatial degrees of freedom of reflective units, thereby improving system capacity, irregular active IRS backscatter communication sends information to users, enhancing the information transmission from the BS to the user. Under the constraints of BS transmit power, backscatter signal-to-interference-plus-noise ratio, and power of irregular active IRS, an optimization problem of maximizing short packet transmission rate is established, and this problem is solved using a DDPG neural network model. Simulation experiments study the model convergence characteristics under different irregularity ratios, the impact of different ratios of reflective units on the system transmission rate under various decoding error probabilities, and the relationship between horizontally moving IRS and the irregularity coefficient. Since the smaller the degree of irregularity, the higher the final converged short packet user transmission rate, and for active IRS, the value changes relatively little during the movement process, because the existence of active IRS itself has an active enhancement effect, so the need to increase the transmission rate by increasing spatial degrees of freedom is very low, thus increasing the system transmission rate.

[0169] Example 2

[0170] Reference Figures 4-6 As an embodiment of the present invention, a short packet transmission method based on irregular IRS-assisted backscatter communication is provided. To verify the beneficial effects of the present invention, a simulation experiment is conducted for scientific demonstration.

[0171] Step 1: Construct a downlink system that supports backscatter short packet communication.

[0172] Consider a communication process from a Base Station (BS) transmitting a signal that is backscattered by an IRS to assist a URLLC user. This process is modeled as a main transmission signal transmitted by the BS and a secondary transmission signal backscattered by the IRS. Assume the BS, IRS, and URLLC user in the communication scenario are in the same two-dimensional coordinate system with coordinates (0,0)m, (100,0)m, and (100,20)m, respectively. Assume the angular parameters of the signal transmission angle and arrival angle follow a random uniform distribution within (0,1). The BS transmission bandwidth is 180kHz, the number of antennas is 4, the Gaussian white noise power spectral density is -170dBm / Hz, the Rician factor is 10, and a half-wavelength ULA is configured. Let d a / λ = 0.5, BS to IRS channel G and IRS to URLLC user channel h r The path loss is 35.6 + 22.0lgDdB, and the direct channel h from the BS to the URLLC user is... d The path loss is 32.6 + 36.7lgD dB, where D represents the horizontal connection distance. Meanwhile, the symbol period L of the backscattered symbol c is set to 50. In fact, the secondary transmission symbol c at the IRS can have its performance adjusted by changing the symbol period L.

[91] Therefore, formula (9) constrains SINR in C2. c =10dB, the transmitted signal power at BS is P T =40dBm. This embodiment considers irregular IRS cells, so for simplicity, p = N / N is used here. S To represent the degree of irregularity, for example, if the number of reflective elements is 60 and the number of grid points on the panel is 120, then the irregularity ratio is p = 0.5. Active IRS panel noise power. The maximum amplification power limit of the active IRS is P. I =30dBm. In order for the active IRS to have a signal amplifier model, the setting of the amplification factor ρ≥1 needs to be very careful, because the incident signal on the IRS panel cannot be amplified indefinitely and needs to be under the constraint of the maximum amplification power of the active IRS.

[0173] Step 2: Build the specific neural network structure of DDPG by setting specific parameters.

[0174] Regarding the parameter settings for the DDPG neural network framework, the number of neurons in the input and output layers of the actor network corresponds to the dimensions of the state space and action space, respectively. The number of neurons in the input and output layers of the critic network corresponds to the cardinality of the state space, action space, and Q-value function, respectively. The two hidden layers in the actor network contain 350 and 250 neurons, respectively, while the first hidden layer in the critic network contains 300 neurons. Since the policy action is also input to the second hidden layer, two temporary layers are used to obtain the corresponding weights and biases, each consisting of 200 neurons. One temporary layer connects to the first hidden layer of the critic network, and the other connects to the action space. All layers in each neural network use the tanh function as the activation function to handle negative values ​​and provide sufficient gradients. For the optimizer selection, the Adam optimizer is used in both the actor and critic evaluation networks, with adaptive learning rates of [missing information]. Where ν Q and ν μ These represent the decay rates of the online actor network and the online critic network, respectively. Other neural network settings include: the size of the experience replay pool. Small batch size The number of episodes is 500, the number of iterations is T = 1000, and the initial learning rate is ζ. Q =ζ μ =0.01, soft update coefficient τ Q =τ μ =0.001, the discount factor for future rewards is γ = 0.99, and the decay rate ν for training network updates is... Q and ν μ The value is 0.0001, and the number of steps for synchronizing the target network and the training network is 1. To enable the agent to explore the communication environment with momentum properties and ensure good correlation in the time series, the OU stochastic process is initialized with random noise.

[0175] Step 3: As described in the invention, the optimization problem is combined with the DDPG network, and the proposed algorithm is experimentally simulated. The conclusions are as follows:

[0176] according to Figure 4 The convergence results show that, under different irregularity ratios, the system model transmission rate convergence diagram obtained using the DDPG framework algorithm has N=20 reflection units and N grid points respectively. S =[20,40,120], corresponding to the irregular proportion p=N / N S= [1, 1 / 2, 1 / 6]. Overall, given the interference power of the reflector panel and the transmit power of the BS, the smaller the p value, the higher the short packet user transmission rate obtained after final convergence, and the greater the improvement effect on the system. Conversely, the larger the p value, the smaller the corresponding improvement effect. When p = 1, the research problem returns to the regular IRS, because the smaller the p value, the higher the spatial degree of freedom of the IRS and the more obvious the spatial diversity advantage. In terms of stages, when the episode < 100, the improvement effect of the small p value is not very good. This is because in the early stage of iteration, it is necessary to continuously try the optimal irregular reflector unit distribution, while the advantage of fixed units is that there is no need to optimize the distribution of reflector elements, so better results can be obtained quickly. At the same time, the system transmission rate of the IRS with p = 1 converges earlier than that of the irregular case. This is because it does not require extra time to explore the optimal reflector unit distribution, while the smaller the p value, the more time is needed to explore the optimal distribution, and the slower the convergence.

[0177] Figure 5 This study investigates the impact of different irregular ratios of IRS reflection units on the system transmission rate for short packet communication users under different decoding error probabilities ε. It shows that when the fixed number of reflection units is N=20, the impact varies with N... S As the number of grid points increases in the interval [20, 120], the system transmission rate also increases. This characteristic is further observed in ε = (10 -5 10 -6 10 -7 All three short packet decoding error probabilities are satisfied. Under the same decoding error probability, the transmission rate of the regular IRS-assisted transmission is lower than that of the irregular case, and the increase is significant. This process is the transformation from a regular arrangement of N=20 to an irregular arrangement of reflective units. Similar to sparse phased array matrices, spatial diversity advantages are achieved by increasing the additional spatial degrees of freedom in the selection of reflective panels, thereby improving system capacity. At p=N / N S =1 to p=N / N S The changes during the p=1 / 6 process are also different. Within the range of p=[1 / 6,1 / 2], the system transmission rate increases rapidly, reaching its lowest point at p=1 / 2. This is because the spatial degrees of freedom are high within this range, allowing the reflection unit to select many grid points, thus resulting in rapid improvement. Within the range of p=[1 / 2,1], the growth rate slows down, the available space decreases, and the spatial diversity advantage becomes less significant. Furthermore, the IRS reflection unit structure with regression rules at p=1 shows no growth trend.

[0178] Figure 6Assuming the BS and user are on the same horizontal plane, and the IRS is moved from the BS to the user with a fixed height, the requirements for the irregularity coefficient p for passive IRS and active IRS under different power enhancement conditions are analyzed. For passive IRS, the irregularity ratio p is largest when deployed directly above the BS or user, and smallest when deployed at the exact center of the horizontal distance. This is because the system transmission rate is highest at the BS or user, and the demand for rate gain is lowest under certain communication quality requirements. Correspondingly, the IRS gain effect is weakest at the center. To achieve a certain communication quality, a higher gain effect is needed, thus requiring a lower p value to allow the reflecting surface to have more spatial degrees of freedom, thereby enhancing the system transmission rate. For active IRS, the overall change in p value during the movement is relatively small. This is because active IRS has a smaller p value due to the higher power enhancement ratio. I The very existence of [the system] has an active enhancement effect, so the need to increase the transmission rate by increasing spatial degrees of freedom is very low. However, the transmission rate of short packets will vary depending on the interference power, and the constraint power P [is also affected]. I The higher the value of the active IRS, the stronger the gain it can provide and the lower the irregularity requirement of the IRS, so the higher the p-value.

[0179] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A short packet transmission method based on irregular IRS-assisted backscatter communication, characterized in that, include: Constructing a downlink short packet communication scenario assisted by irregular active intelligent metasurfaces; Under the constraints of base station transmit power, backscatter signal-to-interference-plus-noise ratio, and power of irregular active intelligent metasurface, an optimization problem model is constructed to maximize the transmission rate of backscatter short packets. The optimization problem model for maximizing the backscatter short packet transmission rate includes: The optimization problem, which aims to jointly optimize the transmission beam at the base station and the reflected beam at the smart metasurface to maximize the short packet transmission rate, is constrained by the transmission power at the base station, the minimum signal-to-noise ratio of backscatter transmission, and the active irregular smart metasurface constraint. The problem can be expressed as follows: in, Let C1 be the transmission beam vector at the base station. This indicates the base station's transmit power. The average signal-to-noise ratio. To constrain the amplification factor, The time indicates that a smart metasurface reflective unit is deployed at the i-th grid point. When , it indicates that no smart metasurface reflection unit is deployed at the i-th grid point, let G represents the channel from the base station to the smart metasurface. When the smart metasurface transmits symbol 1, it is represented as Θ; when it transmits symbol -1, it is represented as -Θ. C3 represents the noise enhancement power; C4 and C5 represent the maximum enhancement power constraint at the smart metasurface; C4 and C5 represent the constraints on the reflective units of the irregular smart metasurface panel; there are a total of [number missing] constraints for the irregular smart metasurface portion. Reflection unit, but the actual number of units used is only N; The optimization problem model is jointly optimized using a deep reinforcement learning model based on a deep deterministic policy gradient. An optimization algorithm is proposed to obtain the maximum neural network feedback in the optimal state, thereby achieving the highest transmission rate.

2. The short packet transmission method based on irregular IRS-assisted backscatter communication as described in claim 1, characterized in that, The downlink short packet communication scenario assisted by the construction of irregular active intelligent metasurfaces includes: The communication scenario includes a base station, an irregular active intelligent metasurface, and a short-packet communication user. The base station has a multi-antenna structure, and the receiver has a single antenna. The system transmits information to the short-packet user via backscatter communication from the irregular active intelligent metasurface. The transmitting base station is equipped with multiple antennas. Low-latency users are all single-antenna users; the intelligent metasurface is composed of... Composed of active reflective units, in which The reflective units are distributed in N s N in each grid space s >N, where the grid constraint with a fixed grid spacing is 1, and the grid spacing between adjacent grid points is set to half the wavelength of the carrier frequency, wherein the irregular active intelligent metasurface can be realized by selecting a certain number of reflective units from all grid points; the intelligent metasurface assists the transmission of the main signal from the base station to the user, and at the same time modulates its own signal onto the incident signal through backscattering technology.

3. The short packet transmission method based on irregular IRS-assisted backscatter communication as described in claim 2, characterized in that, The downlink short packet communication scenario assisted by the construction of irregular active intelligent metasurfaces also includes: This represents the channel from the base station to the smart metasurface. This represents the channel from the base station to the receiving user. This represents the channel from the smart metasurface to the receiving user; each channel consists of two parts: a large-scale fading component and a small-scale fading component. The large-scale fading is distance-dependent and is modeled as follows: in It represents the distance from the signal transmitting end to the receiving end. This represents the path loss per 1m of the reference distance. Indicates the path loss coefficient. This indicates the distance from the base station to the receiving end. This indicates the distance from the base station to the smart metasurface. This represents the distance from the smart metasurface to the receiver; the small-scale fading of each channel follows a Ricean fading channel model, which consists of a line-of-sight (LoS) component and a non-line-of-sight (NLoS) component: in Represents Rice factor, and Representing channels respectively The visual range and the non-visual range; It follows a complex Gaussian distribution with zero mean and unit variance, while the line-of-sight component... The vector representation model is shown below: in , This refers to the space occupied by the antenna. For wavelength, as well as This represents the angle of arrival at the receiver and the angle of transmission at the base station.

4. The short packet transmission method based on irregular IRS-assisted backscatter communication as described in claim 3, characterized in that, The joint optimization includes: Channel from input base station to smart metasurface Channel from smart metasurface to URLLC users Direct channel from base station to URLLC user Output optimal motion space Real-time rewards .

5. The short packet transmission method based on irregular IRS-assisted backscatter communication as described in claim 4, characterized in that, The joint optimization also includes: First, initialize the system; specifically, use the online actor network. Online critic network Target critic network Target actor network And satisfy , Experience replay pool small batch Online learning rate and ; Start the outer loop, which includes: for each Perform the following operations: Get , as well as Initialize the amplification factor Set to 1, initialize with a unit vector. Initialize with the identity matrix , Initialization was handled using the OU noise process. Explore actions and obtain the initial state. .

6. The short packet transmission method based on irregular IRS-assisted backscatter communication as described in claim 5, characterized in that, The joint optimization also includes: Start the inner loop: for each time step Perform the following operations: based on calculate Get instant reward feedback And the new state of the next time slot. The experience gained Stored in the experience replay pool In the middle, from the experience replay pool Medium sampling A random mini-batch of transitions is used to set the target Q-network using the following formula: in , Discount factor, online critic network loss function The presentation format is as follows: Policy gradient for building an online actor network , is represented as: Update the online critic and the online actor network separately, as shown below: in and This represents the network learning rate, used to update the parameters of the online critic network and the online actor network, respectively. The target critic network and the target actor network are soft-updated separately, as follows: in and It also represents the network learning rate, which is used to update the target critic network and the target actor network, respectively; make The inner loop ends, and the outer loop ends.

7. A short packet transmission system based on irregular IRS-assisted backscatter communication, characterized in that, include: The communication scenario building module is used to construct downlink short packet communication scenarios assisted by irregular active intelligent metasurfaces; The optimization problem modeling module is used to construct an optimization problem model that maximizes the backscatter short packet transmission rate under the power constraints of base station transmit power, backscatter signal-to-interference-plus-noise ratio, and irregular active intelligent metasurface. The optimization problem model for maximizing the backscatter short packet transmission rate includes: The optimization problem, which aims to jointly optimize the transmission beam at the base station and the reflected beam at the smart metasurface to maximize the short packet transmission rate, is constrained by the transmission power at the base station, the minimum signal-to-noise ratio of backscatter transmission, and the active irregular smart metasurface constraint. The problem can be expressed as follows: in, Let C1 be the transmission beam vector at the base station. This indicates the base station's transmit power. The average signal-to-noise ratio. To constrain the amplification factor, The time indicates that a smart metasurface reflective unit is deployed at the i-th grid point. When , it indicates that no smart metasurface reflection unit is deployed at the i-th grid point, let G represents the channel from the base station to the smart metasurface. When the smart metasurface transmits symbol 1, it is represented as Θ; when it transmits symbol -1, it is represented as -Θ. C3 represents the noise enhancement power; C4 and C5 represent the maximum enhancement power constraint at the smart metasurface; C4 and C5 represent the constraints on the reflective units of the irregular smart metasurface panel; there are a total of [number missing] constraints for the irregular smart metasurface portion. Reflection unit, but the actual number of units used is only N; The joint optimization module is used to jointly optimize the optimization problem model using a deep reinforcement learning model based on a deep deterministic policy gradient, propose an optimization algorithm, obtain the maximum neural network feedback in the optimal state, and thus obtain the highest transmission rate.

8. A computing device, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, they implement the steps of the short packet transmission method based on irregular IRS-assisted backscatter communication as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the short packet transmission method based on irregular IRS-assisted backscatter communication as described in any one of claims 1 to 6.