Underground Internet of Things architecture time-frequency energy resource allocation system based on intelligent reflecting surface
By introducing centralized uplink data access points, distributed downlink energy access points and intelligent reflection surfaces into the underground Internet of Things, combined with the multi-agent collaborative resource allocation algorithm, the problems of node energy limitation and resource waste in underground tunnels are solved, and efficient resource scheduling and channel optimization are achieved.
Patent Information
- Application Number
- CN202510547799.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-08-01
- Estimated Expiration
- 2045-04-28
AI Technical Summary
The prior art is difficult to effectively solve the problems of node energy limitation and resource shortage in the Internet of Things in underground space, especially the reduction in access point coverage area and waste of spectrum resources caused by narrow and long structures of underground tunnels.
The underground Internet of Things architecture is constructed using centralized uplink data access points, distributed downlink energy access points and intelligent reflection surfaces. The time-spectrum-energy joint allocation algorithm is used to eliminate the multi-path effect and optimize channel transmission performance.
It realizes effective reduction of network energy consumption, improve resource scheduling efficiency, meet the QoS needs of heterogeneous equipment, and optimize channel capacity and resource utilization in underground tunnel scenarios.
Smart Images

Figure CN120417048A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of resource allocation, and particularly relates to a time-frequency-energy resource allocation system for an underground Internet of Things architecture based on intelligent reflecting surfaces. Background Art
[0002] The utilization of urban underground space is increasing rapidly, and corresponding sensing and security challenges are constantly emerging. However, the research on sensing network technology for underground scenarios is still in its initial stage, which restricts the further development and utilization of underground space. With the increase in the number and density of heterogeneous devices, the Internet of Things mainly faces the problems of limited node energy and resource shortage.
[0003] Regarding the problem of limited energy, the mainstream research direction is to minimize the long-term energy consumption of the network, extend the network life cycle by saving energy, and reduce the cost of replacing energy sources such as batteries. However, this method can only delay the replacement of energy sources and cannot truly address the challenge of limited node energy. In recent years, wireless energy transfer technology, as a sustainable power supply method for sensing nodes, has received increasing attention from scholars and become a research hotspot. This technology uses radio frequency signals to achieve highly directional beamforming, and then completes wireless energy transfer, which is particularly suitable for providing efficient power supply for randomly distributed sensing node clusters within a local area. However, the path loss problem leads to large spatial transmission losses and also poses higher requirements for the transmission power.
[0004] Currently, the research on resource allocation problems mainly focuses on the above-ground space. The mainstream research direction is to achieve the maximization of the weighted transmission rate or energy efficiency of the network through the allocation of spectrum or time slot resources, supplemented by the optimization of aspects such as transmission power control and beamforming. At the same time, to address the differences in heterogeneous data transmission, some research also aims to maximize the total energy efficiency on the basis of meeting the requirements of heterogeneous quality of service (QoS). However, different from the circular coverage area of the above-ground space, the long and narrow space structure of underground tunnels makes the coverage area of access points approximately rectangular or a combination of rectangles (at the corners). Therefore, under the same monitoring density, the number of nodes that can be covered by access points is reduced several times, and at the same time, the reduction of adjacent cells also greatly reduces the interference, making the spectrum resources relatively abundant compared to the above-ground space. Therefore, the resource allocation algorithms for the above-ground space are difficult to achieve the best performance in the underground space, resulting in partial waste of resources. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a time-frequency-energy resource allocation system for an underground Internet of Things architecture based on intelligent reflecting surfaces, including:
[0006] A centralized uplink data access point for centrally receiving data uploaded by sensing nodes;
[0007] Multiple distributed downlink energy access points for providing energy supply and control signals to sensing nodes;
[0008] At least one intelligent reflecting surface for establishing an auxiliary channel for sensing nodes at the far end and the outer corner of the corner, eliminating the multipath effect, and improving the channel capacity;
[0009] Multiple sensing nodes, which are divided into different types according to quality of service requirements;
[0010] A resource allocation module for constructing a resource allocation algorithm for multi-agent cooperation, jointly allocating time-spectrum-energy and RIS phase shift configuration based on the resource allocation algorithm for multi-agent cooperation according to different types of quality of service requirements, generating a resource allocation result, and feeding back the resource allocation result to the corresponding access point or intelligent reflecting surface of the policy to optimize the channel transmission performance.
[0011] Preferably, the centralized uplink data access point centrally receives data transmitted by the sensing nodes through the uplink, the distributed downlink energy access point provides wireless energy transmission and control signals to the sensing nodes, and the intelligent reflecting surface enhances the channel quality between the sensing nodes and the centralized uplink data access point by adjusting the phase shift configuration.
[0012] Preferably, the sensing nodes are divided into three types: large data volume type, delay sensitive type, and small data volume type according to quality of service requirements. The sensing nodes obtain energy supply and control signals by accessing the nearest distributed downlink energy access point and upload data to the centralized uplink data access point.
[0013] Preferably, the resource allocation module divides the channel resources into time and spectrum resource blocks according to time slots and spectra, and performs energy transmission in fixed unit time lengths to generate energy resource blocks.
[0014] Preferably, the resource allocation module includes:
[0015] CUAP-agent for solving the optimal time-frequency resource block allocation;
[0016] RIS-agent for solving the optimal phase shift configuration;
[0017] DDAP-agent for solving the optimal energy resource block allocation; and training the multi-agent policy network by the way of centralized training and distributed execution, and setting the global value network.
[0018] Preferably, the signal-to-noise ratios of the uplink and downlink of the sensing nodes are respectively:
[0019]
[0020] where hm,U ,h m,R ,h R,U ,h d,m They are the channel fading matrices from node m to CUAP, node m to RIS, RIS to CUAP, and DDAP-d to node m, α m ∈{0,1} indicates whether node m is assisted by RIS data transmission, is the phase shift diagonal matrix of RIS, N0 is the noise power spectrum density, β(ξ m ) is of type ξ m The basic bandwidth of the data transmission method adopted by node m is spectrum resource block B rb Multiples of B DL is the downlink bandwidth, P m and P DL They are the uplink node transmit power and the downlink DDAP transmit power respectively.
[0021] Preferably, the state of the multi-agent collaborative resource allocation algorithm includes: all channel information, task load, remaining power, and current time slot;
[0022] The observation of CUAP-agent is complete status information;
[0023] The observation of RIS-agent is the information of the sensing nodes assisted by RIS in the state;
[0024] The observation of DDAP-agent is the information of sensing nodes within the coverage range in the state.
[0025] Preferably, the resource allocation module is further used to: calculate channel state information based on a channel model, wherein the channel model includes the superposition of path attenuation and multipath fading, and is used to evaluate the channel quality between the sensing node and the CUAP, and the channel state information is used to guide resource allocation decisions.
[0026] Preferably, the multi-agent collaborative resource allocation algorithm is also used to: calculate the long-term cumulative reward based on the generalized advantage estimate, and calculate the discounted cumulative gain as the fitting target of the value network, realize importance sampling through the probability ratio of the new and old strategies, and limit the strategy update amplitude through strategy clipping.
[0027] Preferably, it also includes: verifying the effectiveness of the framework and algorithm through experiments, evenly distributing various types of sensing nodes in the underground channel, and adjusting the network load by controlling the number of activated nodes to achieve comprehensive performance verification.
[0028] Compared with the prior art, the present invention has the following advantages and technical effects:
[0029] Targeting underground tunnel monitoring scenarios, this paper proposes a frequency-division duplex (FDD)-based IoT framework for RIS (Reference Information System) (RIS). This framework utilizes a collaborative architecture of distributed downlink energy access points and centralized uplink data access points. FDD enables simultaneous downlink energy supply to the sensing node cluster and uplink RIS-assisted data transmission, while also incorporating a short control signal broadcast at the initial stage of energy transmission. Within this framework, the present invention models the resource allocation problem as a joint optimization problem of the time-spectrum-energy three-dimensional resource blocks and RIS phase shift configurations that meet the QoS requirements of heterogeneous devices. A transmission pressure index is defined to reflect the pressure of the remaining load on available resources, with minimizing the weighted sum of energy consumption and transmission pressure as the optimization objective. To efficiently solve this MINLP problem, a multi-agent deep reinforcement learning algorithm based on the PPO algorithm is designed. Considering the absence of competition between agents, a centralized critic network and overall reward-based training for multi-agent collaboration are employed. Experiments demonstrate that the proposed algorithm effectively reduces network energy consumption and improves resource scheduling efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The accompanying drawings, which constitute part of this application, are intended to provide a further understanding of this application. The exemplary embodiments and descriptions of this application are intended to explain this application and do not constitute an improper limitation on this application. In the accompanying drawings:
[0031] Figure 1 This is a schematic diagram of an underground IoT scenario based on RIS according to an embodiment of the present invention;
[0032] Figure 2 is a convergence curve diagram of different strategy algorithms according to an embodiment of the present invention;
[0033] Figure 3 This is a graph showing fluctuations in channel resource efficiency versus the number of active nodes according to an embodiment of the present invention;
[0034] Figure 4 This is a graph showing the fluctuation of energy resource efficiency with the number of active nodes according to an embodiment of the present invention;
[0035] Figure 5 Graph showing the impact of distributed or centralized energy transmission and the presence or absence of RIS on energy consumption according to an embodiment of the present invention. DETAILED DESCRIPTION
[0036] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments in this application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0037] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in certain cases, the steps shown or described can be executed in a different order than here.
[0038] Embodiment 1
[0039] In this embodiment, a time-frequency-energy resource allocation system for an underground Internet of Things architecture based on intelligent reflecting surfaces is provided, including:
[0040] A centralized uplink data access point for centrally receiving data uploaded by sensing nodes;
[0041] Multiple distributed downlink energy access points for providing energy supply and control signals to sensing nodes;
[0042] At least one intelligent reflecting surface for establishing an auxiliary channel for sensing nodes at the far end and the outer corner of the corner, eliminating multipath effects, and improving the channel capacity;
[0043] Multiple sensing nodes, which are divided into different types according to service quality requirements; <{
[0044] A resource allocation module for constructing a resource allocation algorithm for multi-agent collaboration, jointly allocating time-spectrum-energy and RIS phase shift configuration for different types of service quality requirements based on the resource allocation algorithm for multi-agent collaboration, generating a resource allocation result, and feeding back the resource allocation result to the distributed downlink energy access point and the intelligent reflecting surface to optimize the channel transmission performance.
[0045] Scenario description:
[0046] The underground scenario considered in this embodiment consists of two straight tunnel sections and a corner, as shown in Figure 1As shown. The proposed RIS-based underground IoT consists of 1 centralized uplink data access point (CUAP), D uniformly distributed distributed downlink energy access points (DDAPs), two RISs, and M sensing nodes (the set is M = {1, 2,..., m,..., M}). Each access point is configured with L antennas, and each RIS has N reflecting elements. According to the QoS requirements, the sensing nodes are divided into three categories: large volume (LV), delay sensitivity (DS), and small volume (SV). To achieve frequency-division duplex communication, sensing node m will access the nearest DDAP-d to obtain control signals and energy supply, and upload data to the central CUAP. The set of sensing nodes controlled by DDAP-d is denoted as dist(·) represents the distance between two points. The RIS is used to establish an auxiliary channel for sensing nodes at the far end and outside the corner, eliminate the multipath effect, and improve the channel capacity.
[0047] The channel resources are sliced into time and spectrum resource blocks according to time slots and spectra. There are K max spectrum resource blocks and I max time resource blocks, for a total of K max ·I max channel resource blocks. Assume that the energy transmission is carried out over a duration T e , and each time slot has E sub energy resource blocks, for a total of I max ·E sub energy resource blocks.
[0048] Channel model:
[0049] The channel fading is modeled as the superposition of path attenuation and multipath fading. That is:
[0050]
[0051] where λ0 is the unit path attenuation, l is the distance between the transmitter and receiver, Δ is the path attenuation exponent, K is the Rice coefficient, h LOS is the line-of-sight component of the multipath channel, and h NLOS ~CN(0,1) is the multipath component.
[0052] Therefore, the uplink and downlink signal-to-noise ratios of sensing node m are respectively:
[0053]
[0054] where hm,U , h m,R , h R,U , h d,m are the channel fading matrices from node m to CUAP, from node m to RIS, from RIS to CUAP, and from DDAP-d to node m, respectively. α m ∈ {0, 1} indicates whether node m is assisted by RIS for data transmission. is the phase shift diagonal matrix of RIS, N0 is the noise power spectral density, and β(ξ m ) is the basic bandwidth of the data transmission mode adopted by node m of type ξ m which is a multiple of the spectral resource block B rb of the downlink bandwidth B DL where P m and P DL are the uplink node transmission power and the downlink DDAP transmission power, respectively.
[0055] The data transmission rate is calculated by the Shannon formula, i.e.:
[0056] R m-UL = β(ξ m )B rb log2(1 + γ m-UL );
[0057] R m-DL = B DL log2(1 + γ m-DL );
[0058] It should be noted that assuming the downlink control packet size is C0, the downlink data transmission duration is T0 = C0 / min(R m-DL ). Thanks to the extremely small packet size, this process will be completed in a very short time. The number of energy resource blocks in the first time slot is E sch - T0 / T e .
[0059] Problem modeling:
[0060] Since the policy is a time-dependent composite process, the energy of a single transmission may not be used in this instance, and the packet size of the node often remains unchanged. Therefore, the energy efficiency metric is split into a weighted sum of energy consumption and transmission pressure, which are defined as follows:
[0061]
[0062] where EC i represents the energy consumption in time slot i, TP i represents the transmission pressure after time slot i ends, q i,e,m ∈ {0, 1} indicates whether the energy resource block e in time slot i is allocated to node m, pi,k,m ∈ {0, 1} indicates whether the spectral resource block k in time slot i is allocated to node m, C i,m represents the transmission load of node m at the beginning of time slot i, T rb is the length of the time resource block, i.e., the time slot, R m is the uplink data rate, is the delay constraint of node m.
[0063] Therefore, the optimization problem is modeled as the weighted minimization of energy consumption and transmission pressure:
[0064]
[0065] Constraints C1 and C2 are the delay and bit error rate requirements of node m respectively. Constraints C3 and C4 limit that the same resource block can only be allocated to the same device. Constraint C5 is the basic bandwidth constraint required by the transmission mode of the node type. Constraint C6 ensures that the energy of node m at the beginning of time slot i plus the energy of wireless transmission minus the energy consumed by data transmission is greater than the energy threshold Q0 to ensure node wake-up. i,m The energy threshold Q0 for wake-up.
[0066] Multi-agent algorithm based on PPO:
[0067] The optimization problem P is a MINLP problem with a complex variable space and is difficult to solve. To achieve an efficient optimization algorithm, the present invention constructs a multi-agent algorithm based on PPO, sets CUAP-agent to solve the optimal time-frequency resource block allocation, RIS-agent to solve the optimal phase shift configuration, and DDAP-agent to solve the optimal energy resource block allocation, and trains the multi-agent policy network (Actor Networks) in a way of centralized training and distributed execution, with only a global value network (Critic Network) set. The present invention utilizes the optimization idea of deep reinforcement learning for long-term cumulative rewards, and optimizes the strategy within the entire period by optimizing the spectrum, energy allocation, and RIS configuration in each time slot i. The specific process of the algorithm is shown in the pseudocode of Algorithm 1.
[0068] First, set the elements of the Markov process:
[0069] 1) State: Let the state s i = {h m,U , h m,R , h R,U , h D,m , C i,m , Q i,m ,..., i}, including all channel fades, task loads, and remaining battery levels. Among them, the observation of CUAP-agent is The observation of RIS-agent is The observation of DDAP-agent is
[0070] 2) Action: The output action of CUAP-agent is The actions of RIS-agent are: The actions of DDAP-agent are: The total action is
[0071] 3) Reward: To maximize the reward, set the reward value to the negative of the optimization target and add a penalty term Used to penalize the reward for actions that exceed the available time slot.
[0072] r i =-(a·EC i +b·TP i )·ρ(i);
[0073] Next, we consider using the Generalized Advantage Estimator (GAE) to calculate long-term cumulative rewards and the discounted cumulative gain (DCG) as the fitting target for the value network. Based on the PPO algorithm, importance sampling is implemented by using the probability ratio of the new and old policies. Policy pruning is used to limit the policy update range, achieving stable convergence. Finally, parameter updates are performed using gradient descent, as shown in Table 1.
[0074] Table 1
[0075]
[0076] To verify the effectiveness of the designed framework and algorithm, a total of 16 sensing nodes of three types are evenly distributed in an underground tunnel with a length of 40m and a width of 6m. The ratio of the number of nodes of each type is (LV:DS:SV=1:2:7). The network load is adjusted by controlling the number of activated nodes (the node density is between 2.5×10 4 per km 2 to 6.6×10 4 per km 2 In order to achieve comprehensive performance verification, the important parameters related to the experimental scenario are shown in Table 2.
[0077] Table 2
[0078]
[0079] To verify the convergence of the algorithm and the advantages of PPO strategy clipping, the following convergence curve is drawn as follows: Figure 2Among them, the random action strategy is difficult to converge, while the A2C algorithm has large convergence fluctuations, is unstable, and the final average effect is also inferior to the PPO algorithm. The PPO algorithm realizes the importance sampling of empirical data through the ratio of new and old probabilities, and achieves more stable convergence performance through policy clipping.
[0080] Through multiple simulation experiments and calculation of the average results, the following channel resource utilization rate is plotted Figure 3 and energy resource utilization rate Figure 4 .
[0081] According to Figure 3-4 it can be seen that as the number of active nodes increases, the resources allocated for use also gradually increase, but the resource scheduling efficiency remains above 95%. The algorithm proposed in this study can adaptively adjust the resource allocation strategy according to the types of active nodes, etc., and complete the transmission task while saving resources as much as possible.
[0082] To compare the advantages and disadvantages of distributed or centralized energy supply, and the performance improvement brought by RIS assistance, the following change of energy consumption with the number of nodes is plotted Figure 5 .
[0083] From Figure 5 it can be seen that the distributed scheme is more energy-efficient than the centralized scheme under normal load; RIS improves the channel capacity through the auxiliary channel, further reducing the energy consumption; the centralized scheme first reaches the maximum available energy limit of the network. As the number of active nodes increases, the energy-saving advantage of the distributed scheme becomes more significant, and the number of nodes supported at the maximum energy output is more than that of the centralized scheme.
[0084] The above is only a preferred specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A time-frequency energy resource allocation system for an underground Internet of Things architecture based on intelligent reflecting surfaces, characterized in that, including: a centralized uplink data access point for centrally receiving data uploaded by sensing nodes; multiple distributed downlink energy access points for providing energy supply and control signals to sensing nodes; at least one intelligent reflecting surface for establishing an auxiliary channel for sensing nodes at the far end and the outer corner of the corner, eliminating the multipath effect, and improving the channel capacity; multiple sensing nodes, which are divided into different types according to the quality-of-service requirements; a resource allocation module for constructing a resource allocation algorithm for multi-agent cooperation, and based on the resource allocation algorithm for multi-agent cooperation, performing joint allocation of time-spectrum-energy and RIS phase shift configuration according to different types of quality-of-service requirements, generating a resource allocation result, and feeding back the resource allocation result to the corresponding access point or intelligent reflecting surface of the policy to optimize the channel transmission performance.
2. The system according to claim 1, wherein The centralized uplink data access point centrally receives data transmitted by sensing nodes through the uplink, the distributed downlink energy access point provides wireless energy transmission and control signals to the sensing nodes, and the intelligent reflecting surface enhances the channel quality between the sensing nodes and the centralized uplink data access point by adjusting the phase shift configuration.
3. The system according to claim 1, wherein The sensing nodes are divided into three types: large data volume type, delay sensitive type, and small data volume type according to the quality-of-service requirements. The sensing nodes obtain energy supply and control signals by accessing the nearest distributed downlink energy access point and upload the data to the centralized uplink data access point.
4. The system according to claim 1, wherein The resource allocation module divides the channel resources into time and spectrum resource blocks according to time slots and spectra, and performs energy transmission in fixed unit time lengths to generate energy resource blocks.
5. The system according to claim 1, wherein The resource allocation module includes: CUAP-agent for solving the optimal time-frequency resource block allocation; RIS-agent for solving the optimal phase shift configuration; DDAP-agent for solving the optimal energy resource block allocation; and training a multi-agent policy network by the method of centralized training and distributed execution, and setting a global value network.
6. The system according to claim 1, characterized in that, The signal-to-noise ratios of the uplink and downlink of the sensing nodes are respectively: where h m,U , h m,R , h R,U , h d,m are the channel fading matrices from node m to CUAP, from node m to RIS, from RIS to CUAP, and from DDAP-d to node m, respectively. α m ∈ {0, 1} indicates whether node m is assisted by RIS for data transmission. is the phase shift diagonal matrix of RIS, N0 is the noise power spectral density, and β(ξ m ) is that the basic bandwidth of the data transmission method adopted by node m of type ξ m is a multiple of the spectrum resource block B rb . B DL is the downlink bandwidth, and P m and P DL are the uplink node transmission power and the downlink DDAP transmission power, respectively.
7. The system according to claim 1, characterized in that, The state of the resource allocation algorithm for multi-agent cooperation includes: all channel information, task load, remaining power, and current time slot; The observation of CUAP-agent is the complete state information; The observation of RIS-agent is the information of the sensing nodes assisted by RIS in the state; The observation of DDAP-agent is the information of the sensing nodes within the coverage range in the state.
8. The system according to claim 1, wherein The resource allocation module is also used for: calculating channel state information based on a channel model, which includes the superposition of path attenuation and multipath fading, for evaluating the channel quality between the sensing node and the CUAP, and the channel state information is used to guide resource allocation decisions.
9. The system according to claim 1, wherein The resource allocation algorithm for multi-agent cooperation is also used for: calculating the long-term cumulative reward based on the generalized advantage estimation, calculating the discounted cumulative gain as the fitting target of the value network, realizing importance sampling through the probability ratio of the old and new policies, and restricting the policy update amplitude through policy clipping.
10. The system according to claim 1, wherein It also includes: Verify the effectiveness of the framework and algorithm through experiments. Deploy various types of sensing nodes uniformly in the underground passage, and adjust the network load by controlling the number of activated nodes to achieve comprehensive performance verification.
Citation Information
Patent Citations
Communication resource allocation method, system and device based on IRS assistance and medium
CN118201090A
RIS-assisted resource allocation method with V2I link capacity maximization
CN118450510A
Resource control by probability tree convolution production cost valuation by iterative equivalent demand duration curve expansion (AKA. tree convolution)
WO2016040774A1