Intelligent reflecting surface-based underground internet of things architecture time-frequency energy resource allocation system

By introducing centralized uplink data access points, distributed downlink energy access points, and intelligent reflectors into underground IoT, and combining them with a multi-agent collaborative resource allocation algorithm, the problems of limited node energy and scarce resources in underground tunnels were solved, achieving efficient resource allocation and channel optimization.

CN120417048BActive Publication Date: 2026-04-07BEIJING UNIV OF POSTS & TELECOMM
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively address the issues of limited node energy and scarce resources in underground IoT systems. In particular, the narrow and elongated structure of underground tunnels reduces the coverage area of ​​access points while relatively abundant spectrum resources result in poor performance of resource allocation algorithms in underground spaces, leading to resource waste.

Method used

An underground IoT architecture is constructed using centralized uplink data access points, distributed downlink energy access points, and intelligent reflectors. A multi-agent collaborative resource allocation algorithm is used to jointly allocate time, spectrum, and energy, and configure RIS phase shift to optimize channel transmission performance.

Benefits of technology

It effectively reduces network energy consumption, improves resource scheduling efficiency, meets the quality of service requirements of heterogeneous devices, and optimizes channel capacity and resource utilization in underground tunnel scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120417048B_ABST
    Figure CN120417048B_ABST
Patent Text Reader

Abstract

This invention discloses a time-frequency energy resource allocation system based on an intelligent reflective surface for an underground Internet of Things (IoT) architecture. Belonging to the field of resource allocation, this invention designs a frequency-division duplex IoT architecture for underground spaces by jointly using distributed energy access points and a Resource Allocation System (RIS). To achieve efficient utilization of channel and energy resources, based on this IoT architecture, this invention models the resource allocation problem as a joint allocation of time, spectrum, and energy to meet QoS requirements, along with RIS phase shift configuration. This problem is a mixed-integer nonlinear programming problem, difficult to solve in polynomial time. To reduce the complexity of the solution, this invention constructs a multi-agent collaborative resource allocation algorithm based on near-end policy optimization. Overall, the purpose of this invention is to build an efficient monitoring IoT in underground spaces, achieving an energy-saving and highly efficient resource allocation scheme.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of resource allocation, and in particular relates to a time-frequency energy resource allocation system based on an underground Internet of Things architecture with intelligent reflective surfaces. Background Technology

[0002] The utilization of urban underground space is increasing rapidly, and corresponding sensing and security challenges are constantly emerging. However, research on sensing network technologies for underground scenarios is still in its initial stage, limiting the further development and utilization of underground space. With the increase in heterogeneous devices and their density, the Internet of Things (IoT) mainly faces the problems of limited node energy and resource constraints.

[0003] For energy-constrained problems, the mainstream research direction is to minimize the long-term energy consumption of the network, extending its lifespan and reducing the cost of replacing batteries and other energy sources by saving energy. However, this approach can only postpone energy replacement and cannot truly address the challenge of limited node energy. In recent years, wireless power transfer technology, as a sustainable power supply method for sensing nodes, has attracted increasing attention from scholars and become a research hotspot. This technology uses radio frequency signals to achieve highly directional beamforming, thereby completing the wireless power transfer, making it particularly suitable for providing efficient power to clusters of sensing nodes randomly distributed within a local area. However, path loss leads to significant spatial transmission loss and also places higher demands on transmission power.

[0004] Currently, research on resource allocation primarily focuses on terrestrial space. The mainstream research direction aims to maximize the weighted transmission rate or energy efficiency of the network through spectrum or time slot resource allocation, supplemented by optimizations such as transmit power control and beamforming. Simultaneously, to address the differences in heterogeneous data transmission, some research strives to maximize overall energy efficiency while meeting heterogeneous Quality of Service (QoS) requirements. However, unlike the circular coverage area of ​​terrestrial space, the elongated spatial structure of underground tunnels makes the access point coverage area approximately rectangular or a combination of rectangles (at corners). Therefore, at the same monitoring density, the number of nodes that an access point can cover is reduced by several times, and the reduction in adjacent cells significantly reduces interference, making spectrum resources relatively abundant compared to terrestrial space. Consequently, resource allocation algorithms designed for terrestrial space struggle to achieve optimal performance in underground space, resulting in some resource waste. Summary of the Invention

[0005] To address the aforementioned technical problems, this invention provides a time-frequency energy resource allocation system based on an underground Internet of Things architecture using an intelligent reflective surface, comprising:

[0006] A centralized uplink data access point is used to centrally receive data uploaded by sensing nodes;

[0007] Multiple distributed downlink energy access points are used to provide energy supply and control signals to the sensing nodes;

[0008] At least one intelligent reflective surface is used to establish auxiliary channels for sensing nodes at the far end and corner, eliminating multipath effects and improving channel capacity;

[0009] Multiple sensing nodes, which are categorized into different types based on service quality requirements;

[0010] The resource allocation module is used to construct a multi-agent collaborative resource allocation algorithm. Based on the multi-agent collaborative resource allocation algorithm, it performs joint allocation of time, spectrum, and energy and RIS phase shift configuration according to different types of service quality requirements, generates resource allocation results, and feeds the resource allocation results back to the access point or intelligent reflector corresponding to the policy to optimize channel transmission performance.

[0011] Preferably, the centralized uplink data access point centrally receives data transmitted by the sensing nodes through the uplink, the distributed downlink energy access point provides wireless energy transmission and control signals to the sensing nodes, and the intelligent reflector enhances the channel quality between the sensing nodes and the centralized uplink data access point by adjusting the phase shift configuration.

[0012] Preferably, the sensing nodes are divided into three types according to service quality requirements: large data volume type, latency sensitive type, and small data volume type. The sensing nodes obtain energy supply and control signals by accessing the nearest distributed downlink energy access point and upload the data to the centralized uplink data access point.

[0013] Preferably, the resource allocation module divides the channel resources into time and spectrum resource blocks according to time slots and spectrum, and generates energy resource blocks by transmitting energy in fixed unit durations.

[0014] Preferably, the resource allocation module includes:

[0015] CUAP-agent is used to solve for optimal time-frequency resource block allocation;

[0016] RIS-agent is used to solve for the optimal phase shift configuration;

[0017] DDAP-agent is used to solve for the optimal allocation of energy resource blocks; and a multi-agent policy network is trained through centralized training and distributed execution, and a global value network is set.

[0018] Preferably, the uplink and downlink signal-to-noise ratios of the sensing node are respectively:

[0019]

[0020] Where hm,U ,h m,R ,h R,U ,h d,m These are the channel fading matrices from node m to CUAP, node m to RIS, RIS to CUAP, and DDAP-d to node m, respectively, and α. m ∈{0,1} indicates whether node m is assisted by RIS for data transmission. Let N0 be the phase shift diagonal matrix of RIS, and β(ξ) be the noise power spectral density. m ) is of type ξ m The basic bandwidth of the data transmission method used by node m is spectrum resource block B. rb Multiples of B DL For downlink bandwidth, P m and P DL These represent the uplink node transmit power and the downlink DDAP transmit power, respectively.

[0021] Preferably, the state of the multi-agent collaborative resource allocation algorithm includes: all channel information, task load, remaining power, and current time slot;

[0022] The observations from the CUAP-agent provide complete state information;

[0023] The observations of the RIS-agent are information about the sensing nodes that are assisted by RIS in the state;

[0024] The DDAP-agent observes information about the sensing nodes within its coverage area in the state.

[0025] Preferably, the resource allocation module is further configured to: calculate channel state information based on a channel model, wherein the channel model includes the superposition of path attenuation and multipath fading, for evaluating the channel quality between the sensing node and CUAP, and the channel state information is used to guide resource allocation decisions.

[0026] Preferably, the multi-agent collaborative resource allocation algorithm is further used to: calculate the long-term cumulative reward based on generalized advantage estimation, calculate the discount cumulative gain as the fitting target of the value network, achieve importance sampling through the probability ratio of the new and old policies, and limit the policy update magnitude through policy pruning.

[0027] Preferably, it also includes: verifying the effectiveness of the framework and algorithm through experiments, uniformly deploying multiple types of sensing nodes in the underground passage, and adjusting the network load by controlling the number of activated nodes to achieve comprehensive performance verification.

[0028] Compared with the prior art, the present invention has the following advantages and technical effects:

[0029] This invention proposes a frequency division duplex (FDM) IoT framework based on RIS (Resource Allocation System) for underground tunnel monitoring scenarios. This framework employs a collaborative architecture of distributed downlink energy access points and centralized uplink data access points. FDM enables simultaneous frequency division of downlink energy supply and uplink RIS-assisted data transmission for the sensing node cluster, with a short-duplex control signal broadcast added at the initial stage of energy transmission. Within this framework, the resource allocation problem is modeled as a joint optimization problem of time-spectrum-energy three-dimensional resource blocks and RIS phase shift configuration to meet the QoS requirements of heterogeneous devices. A transmission pressure index is defined to reflect the pressure of remaining load on available resources, and minimizing the weighted sum of energy consumption and transmission pressure is used as the optimization objective. To efficiently solve this MINLP problem, a multi-agent deep reinforcement learning algorithm is designed based on the PPO algorithm. Considering the absence of competition between agents, multi-agent collaboration is trained through a centralized Critic network and overall reward. Experiments demonstrate that the proposed algorithm effectively reduces network energy consumption and improves resource scheduling efficiency. Attached Figure Description

[0030] The accompanying drawings, which form part of this application, are used to provide a further understanding of this application. The illustrative embodiments and descriptions of this application are used to explain this application and do not constitute an undue limitation of this application. In the drawings:

[0031] Figure 1 This is a schematic diagram of an underground Internet of Things (IoT) scenario based on RIS, according to an embodiment of the present invention.

[0032] Figure 2 The following are convergence curves of different strategy algorithms in embodiments of the present invention;

[0033] Figure 3 This is a graph showing the fluctuation of channel resource efficiency as a function of the number of active nodes in an embodiment of the present invention.

[0034] Figure 4 This is a graph showing the fluctuation of energy resource efficiency as a function of the number of active nodes in an embodiment of the present invention.

[0035] Figure 5 This diagram illustrates the impact of distributed or centralized energy transmission and the presence or absence of a RIS (Radio Router System) on energy consumption, as shown in the embodiments of the present invention. Detailed Implementation

[0036] It should be noted that, unless otherwise specified, the embodiments and features described in this application can be combined with each other. This application will now be described in detail with reference to the accompanying drawings and embodiments.

[0037] It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases the steps shown or described may be executed in a different order than that shown here.

[0038] Example 1

[0039] This embodiment provides an underground IoT architecture time-frequency energy resource allocation system based on a smart reflector, including:

[0040] A centralized uplink data access point is used to centrally receive data uploaded by sensing nodes;

[0041] Multiple distributed downlink energy access points are used to provide energy supply and control signals to the sensing nodes;

[0042] At least one intelligent reflective surface is used to establish auxiliary channels for sensing nodes at the far end and corner, eliminating multipath effects and improving channel capacity;

[0043] Multiple sensing nodes, which are categorized into different types based on service quality requirements;

[0044] The resource allocation module is used to construct a multi-agent collaborative resource allocation algorithm. Based on the multi-agent collaborative resource allocation algorithm, different types of quality of service requirements are jointly allocated in terms of time, spectrum, and energy, and RIS phase shift configuration is performed to generate resource allocation results. The resource allocation results are then fed back to the distributed downlink energy access point and the intelligent reflector to optimize channel transmission performance.

[0045] Scene description:

[0046] The underground scenario considered in this embodiment consists of two straight tunnel sections and one corner, such as... Figure 1As shown, the proposed RIS-based underground IoT consists of one Centralized Uplink Data Access Point (CUAP), D uniformly distributed Distributed Downlink Energy Access Points (DDAPs), two RISs, and M sensing nodes (set as M = {1, 2, ..., m, ..., M}). Each access point is equipped with L antennas, and each RIS has N reflector elements. Based on QoS requirements, the sensing nodes are divided into three categories: Large Volume (LV), Delay Sensitivity (DS), and Small Volume (SV). To achieve frequency division duplex communication, sensing node m will access the nearest DDAP-d to obtain control signals and energy supply, and upload the data to the central CUAP. The set of sensing nodes controlled by DDAP-d is denoted as . The `dist` function (·) represents the distance between two points. RIS is used to establish auxiliary channels for sensing nodes at remote and corner locations, eliminating multipath effects and improving channel capacity.

[0047] Channel resources are divided into time and spectrum resource blocks according to time slots and spectrum, and K is defined as follows: max Each spectrum resource block and I max There are K time resource blocks in total. max ·I max One channel resource block. Assume energy transmission occurs over time T. e The process is carried out, with a total of E timeslots per time slot. sub One energy resource block, totaling I max ·E sub One energy resource block.

[0048] Channel model:

[0049] Channel fading is modeled as a superposition of path attenuation and multipath fading. That is:

[0050]

[0051] Where λ0 is the unit path attenuation, l is the distance between the transmitter and receiver, Δ is the path attenuation exponent, K is the Rice coefficient, and h LOS h represents the line-of-sight component of the multipath channel. NLOS ~CN(0,1) represents the multipath component.

[0052] Therefore, the uplink and downlink signal-to-noise ratios of sensing node m are respectively:

[0053]

[0054] Where hm,U ,h m,R ,h R,U ,h d,m These are the channel fading matrices from node m to CUAP, node m to RIS, RIS to CUAP, and DDAP-d to node m, respectively. α m ∈{0,1} indicates whether node m is assisted by RIS for data transmission. Let N0 be the phase shift diagonal matrix of RIS, and β(ξ) be the noise power spectral density. m ) is of type ξ m The basic bandwidth of the data transmission method used by node m is spectrum resource block B. rb Multiples of B DL For downlink bandwidth, P m and P DL These represent the uplink node transmit power and the downlink DDAP transmit power, respectively.

[0055] The data transmission rate is calculated using Shannon's formula, namely:

[0056] R m-UL =β(ξ) m B rb log2(1+γ m-UL );

[0057] R m-DL =B DL log2(1+γ m-DL );

[0058] It is worth noting that, assuming the downlink control packet size is C0, the downlink data transmission duration is T0 = C0 / min(R). m-DL Thanks to the extremely small data packets, this process will be completed in a very short time, with the number of energy resource blocks in the first time slot being E. sch -T0 / T e .

[0059] Problem modeling:

[0060] Since the strategy is a time-dependent composite process, the energy from a single transmission may not be used in the current transmission, and the data packet size of a node often remains constant. Therefore, the energy efficiency index is decomposed into a weighted sum of energy consumption and transmission pressure, defined as follows:

[0061]

[0062] EC i TP represents the energy consumption within time slot i. i q represents the transmission pressure after time slot i ends. i,e,m ∈{0,1} indicates whether the energy resource block e within time slot i is allocated to node m, pi,k,m ∈{0,1} indicates whether the spectrum resource block k in time slot i is allocated to node m, C i,m T represents the transmission load of node m at the start of time slot i. rb R is the length of the time resource block, i.e., the time slot. m For uplink data bitrate, Let m be the time delay constraint for node m.

[0063] Therefore, the optimization problem is modeled as a weighted minimization of energy consumption and transmission pressure:

[0064]

[0065] Constraints C1 and C2 represent the latency and bit error rate requirements of node m, respectively. Constraints C3 and C4 restrict the allocation of the same resource block to only one device. Constraint C5 is the basic bandwidth constraint required for the transmission mode of the node type. Constraint C6 ensures that at the beginning of time slot i, the energy of node m plus the energy of wireless transmission minus the energy consumed by data transmission is greater than the energy required to guarantee node Q availability. i,m The energy threshold for awakening is Q0.

[0066] PPO-based multi-agent algorithm:

[0067] The optimization problem P is a MINLP problem with a complex variable space, making it difficult to solve. To achieve an efficient optimization algorithm, this invention constructs a multi-agent algorithm based on PPO. A CUAP agent is set to solve for optimal time-frequency resource block allocation, a RIS agent to solve for optimal phase shift configuration, and a DDAP agent to solve for optimal energy resource block allocation. The multi-agent policy networks are trained through centralized training and distributed execution, while only a global value network (Critic Network) is configured. This invention utilizes the optimization idea of ​​long-term cumulative rewards from deep reinforcement learning. By optimizing the spectrum, energy allocation, and RIS configuration within each time slot i, the policy is optimized for the entire cycle. The specific algorithm flow is shown in the pseudocode of Algorithm 1.

[0068] First, define the elements of a Markov process:

[0069] 1) State: Let state s i ={h m,U ,h m,R ,h R,U ,h D,m C i,m Q i,m ,...,i}, contains all channel fading, task load, and remaining power. The observations of the CUAP-agent are... RIS-agent observations are as follows The observations of DDAP-agent are

[0070] 2) Action: The output action of CUAP-agent is... The actions of RIS-agent are The actions of DDAP-agent are The overall action is

[0071] 3) Reward: To maximize the reward, the reward value is set to a negative number of the optimization target, and a penalty is added. Used to penalize actions that exceed the available time slots.

[0072] r i =-(a·EC) i +b·TP i )·ρ(i);

[0073] Secondly, we consider using the Generalized Advantage Estimator (GAE) to calculate the long-term cumulative reward and calculate the Discounted Cumulative Gain (DCG) as the fitting objective for the value network. Based on the PPO algorithm, we achieve importance sampling by comparing the probability ratio of the new and old policies, and limit the policy update magnitude by policy pruning to achieve stable convergence. Finally, we update the parameters through gradient descent, as shown in Table 1.

[0074] Table 1

[0075]

[0076] To verify the effectiveness of the designed framework and algorithm, 16 sensing nodes of three types were evenly deployed in an underground passage 40m long and 6m wide, with the ratio of each type of node being LV:DS:SV = 1:2:7. The network load was adjusted by controlling the number of activated nodes (node ​​density at 2.5 × 10⁻⁶). 4 units / km 2 Up to 6.6×10 4 units / km 2 To achieve comprehensive performance verification, the key parameters related to the experimental scenario are shown in Table 2.

[0077] Table 2

[0078]

[0079] To verify the convergence of the algorithm and the advantages of the PPO strategy pruning, the following convergence curve was plotted: Figure 2Among these, the random action strategy is difficult to converge, while the A2C algorithm exhibits large fluctuations and instability in convergence, and its final average performance is inferior to that of the PPO algorithm. The PPO algorithm achieves importance sampling of empirical data through the ratio of new to old probabilities and achieves more stable convergence performance through policy pruning.

[0080] Based on multiple simulation experiments and average calculations, the channel resource utilization rate is plotted as follows: Figure 3 and energy resource utilization rate Figure 4 .

[0081] according to Figure 3-4 It is known that as the number of active nodes increases, the resources allocated to them also gradually increase, but the resource scheduling efficiency remains above 95%. The algorithm proposed in this study can adaptively adjust the resource allocation strategy according to the type of active nodes, etc., to complete the transmission task while saving resources as much as possible.

[0082] To compare the advantages and disadvantages of distributed versus centralized energy supply, and the performance improvements brought by RIS assistance, the following graph illustrates the change in energy consumption as a function of the number of nodes. Figure 5 .

[0083] Depend on Figure 5 It can be seen that the distributed solution is more energy-efficient than the centralized solution under normal load; RIS improves channel capacity through auxiliary channels, further reducing energy consumption; the centralized solution first reaches the maximum available energy limit of the network. As the number of active nodes increases, the energy-saving advantage of the distributed solution becomes more significant, supporting more nodes at maximum energy output than the centralized solution.

[0084] The above are merely preferred embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A time-frequency energy resource allocation system based on an underground Internet of Things architecture using intelligent reflective surfaces, characterized in that, include: A centralized uplink data access point is used to centrally receive data uploaded by sensing nodes; Multiple distributed downlink energy access points are used to provide energy supply and control signals to the sensing nodes; At least one intelligent reflective surface is used to establish auxiliary channels for sensing nodes at the far end and corner, eliminating multipath effects and improving channel capacity; Multiple sensing nodes, which are categorized into different types based on service quality requirements; The resource allocation module is used to construct a multi-agent collaborative resource allocation algorithm. Based on the multi-agent collaborative resource allocation algorithm, it performs joint allocation of time-spectrum-energy and RIS phase shift configuration according to different types of service quality requirements, generates resource allocation results, and feeds the resource allocation results back to the access point or intelligent reflector corresponding to the strategy to optimize channel transmission performance. The centralized uplink data access point centrally receives data transmitted by the sensing nodes through the uplink, the distributed downlink energy access point provides wireless energy transmission and control signals to the sensing nodes, and the intelligent reflector enhances the channel quality between the sensing nodes and the centralized uplink data access point by adjusting the phase shift configuration. The sensing nodes are divided into three types according to service quality requirements: large data volume type, latency sensitive type, and small data volume type. The sensing nodes obtain energy supply and control signals by accessing the nearest distributed downlink energy access point and upload the data to the centralized uplink data access point. The uplink and downlink signal-to-noise ratios of the sensing nodes are respectively: ; ; in These are the channel fading matrices from node m to CUAP, node m to RIS, RIS to CUAP, and DDAP-d to node m, respectively. Indicates whether node m is assisted by RIS for data transmission. Here is the phase shift diagonal matrix of RIS. For noise power spectral density, For type The basic bandwidth of the data transmission method used by node m is the spectrum resource block. Multiples of, For downlink bandwidth, These represent the uplink node transmit power and the downlink DDAP transmit power, respectively.

2. The system according to claim 1, characterized in that, The resource allocation module divides the channel resources into time and spectrum resource blocks according to time slots and spectrum, and generates energy resource blocks by transmitting energy in fixed unit durations.

3. The system according to claim 1, characterized in that, The resource allocation module includes: CUAP-agent is used to solve for optimal time-frequency resource block allocation; RIS-agent is used to solve for the optimal phase shift configuration; DDAP-agent is used to solve for the optimal allocation of energy resource blocks; and a multi-agent policy network is trained through centralized training and distributed execution, and a global value network is set.

4. The system according to claim 1, characterized in that, The state of the multi-agent collaborative resource allocation algorithm includes: all channel information, task load, remaining power, and current time slot; The observations from the CUAP-agent provide complete state information; The observations of the RIS-agent are information about the sensing nodes that are assisted by RIS in the state; The DDAP-agent observes information about the sensing nodes within its coverage area in the state.

5. The system according to claim 1, characterized in that, The resource allocation module is also used to: calculate channel state information based on a channel model, wherein the channel model includes the superposition of path attenuation and multipath fading, and is used to evaluate the channel quality between the sensing node and CUAP, and the channel state information is used to guide resource allocation decisions.

6. The system according to claim 1, characterized in that, The multi-agent collaborative resource allocation algorithm is also used to: calculate long-term cumulative rewards based on generalized advantage estimation, calculate the discount cumulative gain as the fitting target of the value network, achieve importance sampling through the probability ratio of the new and old policies, and limit the policy update magnitude through policy pruning.

7. The system according to claim 1, characterized in that, Also includes: The effectiveness of the framework and algorithm was verified through experiments. Various types of sensing nodes were uniformly deployed in the underground passage, and the network load was adjusted by controlling the number of activated nodes to achieve comprehensive performance verification.

Citation Information

Patent Citations

  • Communication resource allocation method, system and device based on IRS assistance and medium

    CN118201090A