Machine learning based spatial resource identification system and method for internet of things networks
The machine learning-based spatial resource identification system addresses spectrum inefficiencies in IoT networks by using reinforcement learning to optimize identification parameters, leading to improved spectrum efficiency and communication performance.
Patent Information
- Application Number
- JP2023198665
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2022-12-08
- Filing Date
- 2023-11-22
- Publication Date
- 2025-05-21
- Estimated Expiration
- 2043-11-22
AI Technical Summary
Existing IoT networks face inefficiencies in spectrum usage due to unused spatial frequency resources, leading to suboptimal performance in device-to-device communication.
A machine learning-based spatial resource identification system and method that utilizes reinforcement learning to optimize identification parameters for IoT devices, enabling them to detect and utilize unused spatial frequency resources for improved spectrum efficiency.
The system enhances spectrum efficiency by identifying and utilizing previously unused resources, thereby improving the performance and energy efficiency of IoT device-to-device communications.
Smart Images

Figure 0007680779000021 
Figure 0007680779000022 
Figure 0007680779000023
Abstract
Description
[Technical field]
[0001] The present invention relates to a machine learning based spatial resource identification system and method for Internet of Things networks. [Background technology]
[0002] The proliferation of Internet of Things (IoT) devices (IoTDs) has caused an exponential increase in wireless data traffic in IoT networks. Existing wireless access methods using Macro Base Stations (MBSs) are unable to accommodate the greatly increased data demands due to poor signal quality received by IoTDs located indoors or at cell boundaries. As a result, the deployment of Femto Base Stations (FBSs) is seen as a practical solution to provide better signal quality to IoTDs. In this architecture, macro cells offload data to Internet of Things devices via the femtocell network. This architecture can facilitate efficient spectrum sharing among Internet of Things devices. In addition, femto base stations, which are inexpensive to build and have flexible configurations, can achieve more efficient spectrum sharing by using spectrum information collected by surrounding Internet of Things devices. FIG. 1 is a diagram showing an example of a conventional Internet of Things network model. As shown in FIG. 1, femto base stations (or edge servers) are deployed in the network topology, and Internet of Things devices transmit and receive data to and from the femto base stations via permitted bands. However, certain spatial frequency resources may remain unused by Internet of Things devices, resulting in spectrum inefficiency. By identifying such unused spatial frequency resources, spectrum efficiency can be improved, and Internet of Things devices can utilize such spatial frequency resources for infrastructure-less device-to-device (D2D) communication. The matters described above as background art are merely intended to enhance understanding of the background of the present invention, and should not be accepted as an admission that they correspond to prior art already known to those skilled in the art. [Prior art documents] [Non-patent literature]
[0003] [Non-Patent Document 1] MK Afzal, YB Zikria, S. Mumtaz, A. Rayes, A. Al-Dulaimi, and M. Guizani, “Unlocking 5G spectrum potential for intelligent IoT: Opportunities, challenges, and solutions,” IEEE Commun. Mag., vol. 56, no. 10, pp. 92-93, Oct. 2018 [Non-Patent Document 2] A. Osseiran et al., “Scenarios for 5G mobile and wireless communications: The vision of the METIS project,” IEEE Commun. Mag., vol. 52, no. 5, pp. 26-35, May 2014 Summary of the Invention [Problem to be solved by the invention]
[0004] Therefore, the technical problem to be solved by the present invention is to provide a machine learning based spatial resource identification system and method for Internet of Things networks that can identify communication resources available for device-to-device communication so as to improve spectrum efficiency. The problems to be solved by the present invention are not limited to those described above, and other problems and advantages of the present invention not mentioned can be understood from the following description and will become more apparent from the examples of the present invention. In addition, a person skilled in the art to which the present invention pertains will easily understand that the problems and advantages of the present invention can be realized by the means and combinations thereof set forth in the claims. [Means for solving the problem]
[0005] As a means for solving the above technical problem, the present invention provides a spatial resource identification system for an Internet of Things network, including: a plurality of Internet of Things devices that use an identification parameter to identify available spatial resources and provide corresponding spectrum identification information as a result; and a wireless resource harvesting edge that receives the spectrum identification information, optimizes the identification parameter based on the received spectrum identification information, and provides the optimized identification parameter to the Internet of Things devices. In one embodiment of the present invention, the Internet of Things device can use the available spatial resources to perform device-to-device communications with other Internet of Things devices. In one embodiment of the present invention, the Internet of Things device can detect surrounding Internet of Things devices using an energy sensing method based on the signal-to-noise ratio of a received signal. In one embodiment of the present invention, the radio resource harvesting edge may be a learning agent that executes a reinforcement learning algorithm to optimize the discrimination parameter. In one embodiment of the present invention, the state of the reinforcement learning algorithm may be a beam set consisting of beams of the Internet of Things device, and the action of the reinforcement learning algorithm may be a subset of the beam set consisting of beams selected from the beams of the Internet of Things device. In one embodiment of the present invention, the compensation of the reinforcement learning algorithm may be defined as the probability of searching for spectrum resources relative to the energy consumed to identify the Internet of Things device. In one embodiment of the present invention, the radio resource harvesting edge may apply a c-greedy algorithm for the action of the reinforcement learning algorithm. In one embodiment of the present invention, the radio resource harvesting edge can collect the spectrum identification information from the Internet of Things device, execute the reinforcement learning algorithm to calculate compensation for optimal beam set determination, share identification parameters derived based on previously collected spectrum identification information and previously derived identification parameters with the Internet of Things device, and retrain and search for optimal identification parameters based on results reported from the Internet of Things device. As another means for solving the technical problem, the present invention provides a spatial resource identification method for an Internet of Things network, including: an identification step in which each of a plurality of Internet of Things devices searches for unused spatial resources using an identification parameter provided by a wireless resource harvesting edge, and the wireless resource harvesting edge searches for the optimal identification parameter through reinforcement learning; a reporting step in which the plurality of Internet of Things devices send an identification result to the wireless resource harvesting edge through a control channel, and the wireless resource harvesting edge receives the identification result and sends the identification parameter searched through reinforcement learning of the previous identification step to the Internet of Things device; and a communication step in which the plurality of Internet of Things devices perform device-to-device communication through surrounding unused spectrum resources, and the wireless resource harvesting edge updates the identification result received in the reporting step and starts searching for an identification parameter using reinforcement learning. Effect of the Invention
[0006] The spatial resource identification system and method for an Internet of Things network applies machine learning, particularly reinforcement learning, to identify unused spectrum space not used by Internet of Things devices with less energy, thereby improving overall network resource efficiency and thereby improving performance. The effects obtained by the present invention are not limited to those mentioned above, and other effects not mentioned will be clearly understood by those having ordinary skill in the art to which the present invention pertains from the following description. [Brief description of the drawings]
[0007] The drawings attached to this specification illustrate preferred embodiments of the present invention and, together with the detailed description of the invention described below, serve to facilitate understanding of the technical concepts of the present invention. Therefore, the present invention should not be interpreted as being limited solely to the matters described in the following drawings. [Figure 1]FIG. 1 is a diagram illustrating an example of a conventional Internet of Things network model. [Diagram 2] 1 is a diagram illustrating a network model to which a spatial resource identification system according to an embodiment of the present invention is applied and its operation. [Diagram 3] 1 is a graph showing experimental results for evaluating the applicability of a spatial resource identification system based on reinforcement learning according to an embodiment of the present invention. [Figure 4] FIG. 2 illustrates a workflow of a reinforcement learning-based spatial resource identification system according to an embodiment of the present invention. [Diagram 5] 1 is a graph comparing simulation results of energy consumption versus average operation time of an Internet of Things device of a spatial resource identification system based on reinforcement learning according to an embodiment of the present invention and a conventional spatial resource identification system. [Figure 6] 1 is a flowchart illustrating an optimal communication execution technique during outages on a wireless resource harvesting edge according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0008] The specific structural or functional description of the embodiments described below is disclosed for illustrative purposes only, and may be modified and implemented in various forms. Therefore, the embodiments are not limited to the specific disclosed forms, and the scope of the present specification includes modifications, equivalents, or alternatives within the technical spirit. Although terms such as first or second may be used to describe various components, such terms should be construed only for the purpose of distinguishing one component from another component, for example, a first component may be termed a second component, and similarly, a second component may be termed a first component. When a component is referred to as being "coupled" to another component, it should be understood that the component may be directly coupled or connected to the other component, but there may also be other components in between. The singular term includes the plural term unless the context clearly dictates otherwise. In this specification, the terms "comprise" or "have" and the like are intended to specify the presence of a stated feature, number, step, operation, component, part, or combination thereof, and should be understood as not precluding the possibility of the presence or addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof. Unless otherwise defined, all terms used herein, including technical or scientific terms, have the same meaning as commonly understood by a person of ordinary skill in the art. Terms as defined in commonly used dictionaries should be interpreted to have a meaning consistent with the meaning they have in the context of the relevant art, and should not be interpreted as having an idealized or overly formal meaning unless expressly defined herein.
[0009] First, each element of the spatial resource identification system according to the embodiment of the present invention and the basic concepts applied thereto will be described in detail. FIG. 2 is a diagram illustrating a network model to which the spatial resource identification system according to the embodiment of the present invention is applied and its operation. As shown in FIG. 2, a network model to which the spatial resource identification system according to the embodiment of the present invention is applied may include an Internet of Things device (IoTD) 11 and a Radio Resource Harvesting Edge (RRHE) 20. In addition, in the spatial resource identification system according to the embodiment of the present invention, the Internet of Things device may operate in three stages: an identification stage (S1), a reporting stage (S2), and a transmission stage (S3). In each stage, the Internet of Things device 11 may perform the following operations. In FIG. 2, reference numeral '10' denotes an Internet of Things device that communicates with a base station, and since the embodiment of the present invention relates to a spatial resource identification technique that enables device-to-device communication between the Internet of Things, in the description of the embodiment of the present invention, the Internet of Things device may be understood as an Internet of Things device indicated by reference numeral '11' that undergoes a spectrum identification process.
[0010] 1) Identification step (S1): Each Internet of Things device can sense signals within the sensing range of the Internet of Things device based on an energy sensing method, and identify available space resources. The identification method applied in the embodiment of the present invention can use optimized identification parameters and store the positions of the sensed Internet of Things devices in a local connection matrix.
[0011] 2) Reporting Phase (S2): Each Internet of Things device can send a local connectivity matrix including location information of the Internet of Things devices detected in the identification phase (S1) to the Radio Resource Harvesting Edge (RRHE). The report frame can be transmitted via a basic control channel or in a multi-hop manner via wireless backhaul routing. The Radio Resource Harvesting Edge (RRHE) can identify optimal identification parameters based on the transmitted information and transmit this information to each Internet of Things device in a subsequent reporting phase.
[0012] 3) Transmission stage (S3): In this stage (S3), Internet of Things devices can communicate by utilizing unused spatial frequency resources. FIG. 1 also shows a timeline applied to an embodiment of the present invention. In an embodiment of the present invention, each Internet of Things device can identify unused spectrum resources using an identification parameter in each identification step (S1). After the identification step (S1), each Internet of Things device can transmit a report frame including an identification result to a radio resource harvesting edge (RRHE) and receive a command frame from the RRHE. After receiving the command frame, each Internet of Things device can block a beam recorded in the received command frame. In a transmission step (S3), the Internet of Things device can transmit and receive data frames to other Internet of Things devices using the unblocked beam. In this step (S3), the Internet of Things device can operate in a contention-based manner. More detailed related communication techniques will be described later.
[0013] Internet of Things devices that perform device-to-device (D2D) communication are equipped with directional antennas and can perform directional communication. Directional antennas in the mmWave band can be classified into 1) switched-beam antennas and 2) beam-steering antennas. Switched-beam antennas are designed to cover a certain area per fixed beam, and one beam can be activated to perform communication. In beam-steering antennas, the main beam can be controlled by a phase shifter to the desired direction to transmit and receive information. The switched beam antenna has the advantage of being easy to implement and inexpensive, but has the disadvantage of weakening the signal strength when switching between beams. The beam steering antenna has the advantage of implementing high signal quality through sophisticated control, but is expensive and complicated to implement. Since Internet of Things systems tend to have limited available energy and computer performance, it can be assumed that a switched beam antenna is used in the embodiments of the present invention. It can also be assumed that a wireless Internet device including a radio resource harvesting edge (RRHE) is equipped with a switched beam array antenna having M beam patterns, and that the beam patterns do not ideally overlap. During transmission, only one direction of each sector can be activated to transmit a signal, and the remaining sectors can be blocked. During reception, many sectors can be activated simultaneously or only a specific direction can be activated. It can also be assumed that an antenna controller is used to track the direction in which the maximum signal power is received.
[0014] A radio resource harvesting edge (RRHE) is an element that collects identification results from Internet of Things devices. Based on the identification results, the radio resource harvesting edge can determine optimal identification parameters for each Internet of Things device. In addition, the radio resource harvesting edge can transmit the identified spatial frequency information to the Internet of Things device for device-to-device communication. The radio resource harvesting edge can include a macro base station (MBS), a femto base station (FBS), a WiFi access point (AP), or any dedicated RRHE form or combination thereof. Although a typical FBS / MBS knows frequency information of a cell area, it cannot know local information of an Internet of Things device. Therefore, the radio resource harvesting edge can maximize frequency resource efficiency by allocating spatial frequency resources to the Internet of Things device based on local identification information.
[0015] An Internet of Things device can use C data channels and a non-orthogonal multiple access (NOMA) wireless network. The NOMA method is realized by combining orthogonal frequency division multiple access and multi-carrier code division multiple access. As known in these techniques, the embodiment of the present invention assumes a single physical data channel having S subcarriers. The data channel can be divided into two subcarrier groups (SubCarrier Groups: SCGs) as follows:
[0016] 1) SCG1(S 1 ): The minimum number of subcarriers for data transmission. Such subcarriers occupy only a small portion of the data channel bandwidth. 2) SCG2(S 2 ): Data channel S 1 The remaining subcarrier set excluding In an embodiment of the present invention, the entire subcarriers can be allocated to each beam of an Internet of Things device i. i j denotes the data channel in the j-th beam direction of the Internet of Things device i. C i j can be interpreted as the geographical transmission and reception coverage area of the Internet of Things device i when using the j-th beam. i 1、j and S i 2、j Let SCG1 and SCG2 denote the j-th beam of the Internet of Things device i, respectively. An embodiment of the present invention can be designed to have a wider spectrum for S2 than for S1, and S 1 and S. 2 It can be assumed that the and are sufficiently separated that the interference between them can be neglected.
[0017] In addition, in the embodiment of the present invention, all S 1The control channel can be regarded as an underlying control channel that utilizes the above-mentioned. The implementation of the basic control channel for a cognitive radio network (CRN) has already been verified in the art. Each Internet of Things device can report the identification result (presence and location of an IoTD) to the radio resource harvesting edge via the control channel. Similarly, the radio resource harvesting edge can use the reported information to calculate optimal identification parameters for each Internet of Things device and propagate the identification parameters to all Internet of Things devices via the control channel. Several classical spectrum identification techniques have been proposed, including matched filters, feature sensing, and energy sensing, and in an embodiment of the present invention, an Internet of Things device can utilize energy sensing techniques to determine the presence or absence of an Internet of Things device based on the amount of energy received. The received signal can be integrated every observation period, and finally, the output of the integrator can be divided by the noise power [i.e., signal-to-noise ratio (SNR)] and compared with a certain critical value (or the sensing sensitivity of the wireless Internet device) to confirm the presence of the wireless Internet device. If the signal-to-noise ratio is greater than the sensing sensitivity of the Internet of Things device, it can be determined that an Internet of Things device is present on the identified channel.
[0018] In the following, the main operations of the spatial resource identification system according to the embodiment of the present invention will be described in detail. Directional Spectrum Discrimination / Harvesting In the cooperative directional identification / harvesting method, the Internet of Things devices share identification regions with each other, so that overlapping regions are not detected. Also, the wireless resource harvesting edge can form multiple clusters. Each Internet of Things device can identify a specific direction assigned to it using an identification beam. Such identification beams are assigned to the Internet of Things devices by the wireless resource harvesting edge, which can attempt to maximize the probability of detecting the Internet of Things device while reducing the overall identification overhead. To improve the detection probability, the Internet of Things device can use as many beams as possible for identification. However, such an approach may result in various side effects such as considerable energy consumption by the Internet of Things device, in addition to the time and network resources consumed for identification. The following are the main overhead issues caused by the identification process. 1) Energy Overhead Due to Identification: Internet of Things devices that perform identification consume additional energy. 2) Energy overhead due to reporting: An Internet of Things device that transmits its identification result to the wireless resource harvesting edge consumes energy to transmit the report frame. Therefore, optimization techniques must identify identification parameters that can sense Internet of Things devices while reducing the identification overhead.
[0019] Identification parameter optimization In order to maximize the efficiency of the spectrum identification process, an embodiment of the present invention uses an objective function to select a beam set (B i ) can be optimized. i "Bi = [b i 0 , b i 1 , b i 3 , ... and b i M-1]T, where b i j is 1, otherwise b i j is 0. Here, the objective function can be expressed as the following Equation 1.
[0020]
number
[0021] where P(B) and O(B) respectively denote the Internet of Things device detection probability and identification overhead for B, α denotes a weighting factor, and B denotes the beam set for every Internet of Things device i. Here, the Internet of Things device detection probability P can be modeled as Equation 2:
[0022]
number
[0023] Here, N is the number of Internet of Things devices, and P(B i ) denotes the detection probability of the i-th Internet of Things device in terms of the identified beam, where P i (B i ) can be given as Equation 3:
[0024]
number
[0025] Conventionally, identification overhead is modeled in terms of time overhead for identification, but the embodiment of the present invention can model identification overhead (O) based on the energy consumed for identification. In the embodiment of the present invention, identification overhead O can be defined as the following Equation 4.
[0026]
number
[0027] Here, O s (B) shows the overhead due to spectrum identification, and O r (B) shows the overhead due to transmitting the identification result and receiving the identification parameters from the Radio Resource Harvesting Edge (RRHE). The wireless Internet device consumes energy while identifying the channel. Let “ρ” denote the energy consumption per unit time, then O s (B) can be expressed as the following equation 5.
[0028]
number
[0029] Here, t s and |B i | denotes the discrimination time and the number of beams selected for spectrum discrimination, respectively. Similarly, the Internet of Things device also consumes energy while transmitting the identification result and receiving the identification parameters, resulting in reporting overhead O r (B) can be determined according to the following equation 6.
[0030]
number
[0031] Here, L r indicates the length of the reporting phase. In the past, in order to improve identification accuracy and minimize overlapping identification areas to reduce overhead, a method was applied in which the identification beam with the highest signal-to-noise ratio (SNR) value was used. However, in this conventional technique, since a beam with a good channel condition is always selected, there was a problem that energy consumption imbalance occurred between Internet of Things devices, and since identification overhead was modeled as a wasted time opportunity, the energy consumed for identification was relatively high. In order to solve such a conventional problem, an embodiment of the present invention proposes a beam selection algorithm based on a machine learning technique.
[0032] Optimal beam selection based on reinforcement learning (RL) A standard RL model includes a finite set of possible states of the environment S = {s1, s2, ..., sn}, a set of possible actions of the learning agent A = {a1, a2, ..., am}, a scalar reinforcement signal r, and an agent policy π. At each time step, the agent perceives a state of the environment s ∈ S and selects an action a ∈ A according to the current policy π. Time is represented as a sequence of time steps t = 0, 1, ... At each time step, the controller observes the current state of the system and selects an action, which shifts the environment to a new state s' ∈ S and immediately delivers a compensation reinforcement signal c t The new state and reinforcement signal are presented to the learning agent, which updates its policy and the next iteration round begins. The goal of the learning agent is to find an optimal policy for each state, π, that minimizes the total expected discount compensation over an infinite time horizon. * Such compensation can be defined by the following Equation 7:
[0033]
number
[0034] where E denotes the operator's expectation and γ∈(0,1) is the discount factor. An RL algorithm is considered to converge when the learning curve flattens out and does not increase any more. It has been proven that Q-learning converges to an optimal solution. Therefore, the optimality condition can be defined as the following Equation 8.
[0035]
number
[0036] Here, t denotes the iteration step and e denotes the small size critical value. According to Bellman's optimality criterion, the optimal policy π satisfies the following equation 9.
[0037]
number
[0038] Here, C(s, a) denotes the expected cost C(s, a) = E{c(s, a)}, and P s、s’ denotes the transition probability for a change from s to s'. Given the optimal value function, the optimal policy can be specified as shown in Equation 10.
[0039]
number
[0040] For each learning agent i, the evaluation function, denoted Q(s, a), can be defined as the expected discounted reinforcement credited to the optimally chosen action after taking action a in state s, as shown in Equation 11:
[0041]
number
[0042] For each learning agent i, Q(s, a) can be rewritten as:
[0043]
number
[0044] To apply Bellman's criterion, Q * We must find an intermediate minimal value of Q(s, a), denoted by (s, a), where an intermediate criterion function for all possible successor state-action pairs is minimized and the optimal action is taken for each successor state. Q * (s, a) is given by the following equation 13.
[0045]
number
[0046] And the operation a for the current state s * In other words, π * Therefore, Q * (s, a * ) is minimal and can be expressed as follows:
[0047]
number
[0048] Using the information (s, a, s', a') available in the Q-learning process, an attempt can be made to find Q(s, a) in a recursive manner, where s and s' are the states at times t and t+1, respectively, and a and a' are the actions taken at times t and t+1, respectively. The Q-learning rule for updating the Q-value for learning agent i is given by Equation 15 below.
[0049]
number
[0050] Here, α denotes the learning rate. In an embodiment of the present invention, it is possible to try to find an optimal beam set for all Internet of Things devices under different environmental conditions so that the objective function is minimized. The reason for applying reinforcement learning to select the optimal beam set is as follows. 1) The advantage of reinforcement learning is that the optimal results increase over time due to compensation-based learning. In a network environment, Internet of Things devices are distributed and can be considered as opportunistically sensing unused spectrum resources for device-to-device communication. Thus, a fixed Internet of Things device can sense a large amount of unused spectrum at a lower cost over time. 2) The network topology may change frequently. The state of an Internet of Things device may change at any time. In some cases, the battery may be depleted and the Internet of Things device may be turned off, or a new Internet of Things device may join the network. Therefore, in a topology where network conditions change frequently, reinforcement learning techniques that produce optimal results with relatively few operations may be appropriate. In an embodiment of the present invention, a radio resource harvesting edge (RRHE) can be a learning agent, and an Internet of Things device and a set of discriminatory beams of the Internet of Things device can be an environment of the learning agent. Thus, in an embodiment of the present invention, a basic reinforcement learning element can be defined as follows: 1) State: The selection of the state space is a fundamental step in Q-learning. The selected state variables must be free of side effects and contain knowable features. In the embodiment of the present invention, a state is defined as the set of beams that each Internet of Things device utilizes for identification. MIf beam combinations are used, the number of states of each Internet of Things device is 2 M and the state of the Internet of Things device i (s i ) can be defined as follows:
[0051]
number
[0052] Here, if the j-th beam of the Internet of Things device i is used for identification, then b i j is 1, otherwise b i j is 0. 2) Action: A set of possible actions can be determined for each Internet of Things device based on the selected beam. The action of Internet of Things device i (a i ) is expressed as the following equation 17: i ) can be defined as a subset of
[0053]
number
[0054] Where: TIFF0007680779000018.tif1811 indicates the members of the selected beam set, and S indicates the number of selected beams. 3) Reward: In an embodiment of the present invention, in RL-based Q-learning, each state is defined as a set of beams that each Internet of Things device utilizes for identification, and an action is defined as a set of beams to be activated for all node identification. When determining new beam sets for all Internet of Things devices, the overall network reward is defined as the probability of searching for spectrum resources relative to the energy consumed for identification. That is, the reward function C(s, a) can be designed using the objective function (Equation 1). Therefore, the reward function C(s, a) can be defined as the following Equation 18.
[0055]
number
[0056] In the classification algorithm according to the embodiment of the present invention, the radio resource harvesting edge (RRHE) determines the classification parameters of all the Internet of Things devices. Since the channel environment, location, etc. are different for each Internet of Things device, the radio resource harvesting edge must learn individually for each Internet of Things device. The following algorithm 1 determines a beam set B based on Q-learning for spectrum classification. i Describes the selection algorithm.
[0057] [Table 1]
[0058] For beam selection (action) by each Internet of Things device, the Radio Resource Harvesting Edge can utilize a c-greedy algorithm, where "c" is an arbitrary factor used to search for optimal values to avoid local minima. The radio resource harvesting edge may operate as follows according to a reinforcement learning (RL) technique in accordance with an embodiment of the present invention. 1) The wireless resource harvesting edge collects the spectrum identification information of the Internet of Things devices in the reporting stage. 2) The radio resource harvesting edge executes the reinforcement learning algorithm until the next reporting stage and calculates the compensation for the optimal beam set decision. 3) In the next reporting stage, the radio resource harvesting edge shares the derived identification parameters with the Internet of Things device based on previously acquired values (such as previously reported spectrum identification information and previously derived identification parameters), and retrains and searches for optimal values based on the reported results.
[0059] In an embodiment of the present invention, the wireless resource harvesting edge may include a learning agent that must manage the identification parameters of all the Internet of Things devices. Therefore, the algorithm complexity may vary depending on the number of Internet of Things devices and the number of antennas. The embodiment of the present invention may analyze the algorithm complexity and search for an optimal identification beam set based on reinforcement learning. In an embodiment of the present invention, each Internet of Things device may have a total of 2 M can have two states M There can be candidate actions. Therefore, the proposed algorithm to search for the optimal beam set based on reinforcement learning takes O(N 2 M ) can have a worst-case time complexity of 100 ms. In the algorithm for searching the optimal beam set, the wireless resource harvesting edge searches for the optimal beam set for all the Internet of Things devices based on reinforcement learning, and can inform each Internet of Things device of the optimal beam set in the next reporting stage. Therefore, the calculation must be completed within the interval of the reporting stage. We will explain an example of implementing wireless resource harvesting edge using Raspberry Pi 4B as a representative Internet of Things device. In this example, the CPU generates a clock signal at about 1.5 GHz. Assuming that about 100 clock cycles are required to calculate the compensation, the 7 If M=16 (the number of antennas in the Internet of Things device=16) and the identification period is 1 second, then 228 Internet of Things devices (228×2 16 ≒1.5×10 7 ) can be searched for the optimal discrimination beam. Therefore, the discrimination period is set to T s So, 228×T s It is possible to search for an identifying beam for an Internet of Things device.
[0060] In order to evaluate the applicability of the spatial resource identification technique based on the reinforcement learning described above, the inventors of the present invention constructed a simple Internet of Things system and measured the calculation time. In this system, a Raspberry Pi 3 model plays the role of a wireless resource harvesting edge, and reinforcement learning is used to search for the optimal beam of 200 surrounding Internet of Things devices. FIG. 3 is a graph showing the results of an experiment for evaluating the applicability of a spatial resource identification system based on reinforcement learning according to an embodiment of the present invention. As shown in Figure 3, the experimental results show that compensation converges between about 150 and 200 epochs. Since 20 epochs consume about 1 second, it was analyzed that it takes about 7.5 seconds in the actual experimental environment. Therefore, to guarantee accurate sensing results in this system, optimization must be completed within 7.5 seconds. In other words, the sum of the identification period (Ts) and communication period (Tc) must be longer than 7.5 seconds.
[0061] Reinforcement learning based classification technique workflow 4 is a diagram showing the workflow of the spatial resource identification system based on reinforcement learning according to an embodiment of the present invention. As shown in FIG. 4 and described above, the spatial resource identification technique according to an embodiment of the present invention can be composed of three stages (P1, P2, P3). In the identification phase (P1), each Internet of Things Device (IoTD) can search for unused spatial resources using the identification parameters selected by the Radio Resource Harvesting Edge (RRHE). At the same time, the Radio Resource Harvesting Edge (RRHE) can search for optimal identification parameters by reinforcement learning (RL). In the reporting stage (P2), each Internet of Things device (IoTD) can send the identification result to the Radio Resource Harvesting Edge (RRHE) via a control channel. The Radio Resource Harvesting Edge (RRHE) can receive the identification result and send the identification parameters found in the previous reinforcement learning (RL) stage to the Internet of Things device (IoTD). In the communication phase (P3), each Internet of Things device (IoTD) performs device-to-device (D2D) communication via the surrounding unused spectrum resources, and the radio resource harvesting edge (RRHE) updates new data (identification results) and begins to identify identification parameters using reinforcement learning. In the communication phase (P3), the radio resource harvesting edge (RRHE) can perform the same operation as in the identification phase (P1). The process performed by the radio resource harvesting edge (RRHE) in the communication phase (P3) and the identification phase (P1) can be said to be a learning and optimization process.
[0062] FIG. 5 is a graph comparing simulation results of energy consumption versus average operation time of an Internet of Things device for a spatial resource identification system based on reinforcement learning according to an embodiment of the present invention and a conventional spatial resource identification system. Referring to FIG. 5, it can be seen that in both the embodiment of the present invention and the conventional technology, the longer the operation time of the Internet of Things device, the more frequently the spectrum identification is performed, and therefore the energy consumption increases. In particular, it was observed that the conventional omni-directional identification technique (black solid line) consumes more energy for identification than the directional identification technique. This is because the identification range of the omni-directional identification technique is smaller than that of the directional identification method. This means that the omni-directional identification method requires more energy consumption to increase the sensing distance compared to the directional identification method. It can also be seen that the spectrum identification technique based on reinforcement learning according to the embodiment of the present invention (blue dotted line) shows a performance improvement of about 18% compared to the conventional directional identification method (red solid line). This is because the identification work of the conventional directional identification method is performed by a specific Internet of Things device, whereas the identification work is uniformly distributed in the embodiment of the present invention. Embodiments of the present invention provide techniques for performing optimal communication between Internet of Things devices 10, 11 via spatial resource discrimination when the wireless resource collection edge 20 is inoperable due to a power interruption or external attack.
[0063] FIG. 6 is a flowchart illustrating a technique for optimal communication execution during unavailable times in a wireless resource collection edge according to an embodiment of the present invention. First, when the operation of the wireless resource collection edge becomes impossible (S11), the Internet of Things devices 10, 11 broadcast information about their own positions using LoRa (Long Range). Then, the Internet of Things devices 10, 11 may perform grouping based on this position information and determine a leader in the group (S12). This grouping and leader determination can be performed through a preset algorithm based on the exchanged position information. Then, within each group, wireless resource allocation can be performed by repeated information exchange between a leader Internet of Things device (hereinafter referred to as "leader") and the remaining Internet of Things devices (hereinafter referred to as "followers") (S13). Here, the follower can apply a preset algorithm to execute a beam selection strategy that can maximize the sensing probability with the minimum sensing power. By executing this strategy using the information on the optimal beam, which is a control variable received from the leader, the follower generates information on the beam and information on the sensing power threshold of the beam and transmits it to the leader. The leader can apply a pre-configured algorithm to apply a strategy to prevent overlap between the sensed beams. The leader can utilize the information received from the followers to generate and transmit to each follower a control variable corresponding to information regarding an optimal beam that can prevent overlap between the beams. After that, the follower and the leader repeatedly transmit and receive information implementing the strategy described above. The leader ends the iteration when it reaches an inflection point where the preset objective function calculated using the received information decreases and increases. The leader can determine the beam with the received information at the end as the beam to apply to communication and turn off the remaining beams (S14). The technique shown in Figure 6 is based on Stackelberg game theory and achieves radio resource allocation by finding a point that is a Stackelberg-Nash equilibrium (SNE) and selecting the optimal beam. [Explanation of symbols]
[0064] 10 Internet of Things devices (devices that communicate with base stations) 11 Internet of Things Devices (Spectrum Identification Devices) 20 Radio Resource Harvesting Edge P1 Identification stage P2 Reporting stage P3 Transmission Stage
Claims
1. A plurality of Internet of Things devices that identify available spatial resources in their vicinity and provide spectrum identification information corresponding to the identified resources; a wireless resource harvesting edge that receives the spectrum identification information, and derives optimal communication-enabled spatial resources based on the received spectrum identification information and provides the optimal communication-enabled spatial resources to the Internet of Things device; the wireless resource harvesting edge is a learning agent that executes a reinforcement learning algorithm to derive the optimal communication-enabled spatial resource; a state of the reinforcement learning algorithm is a beam set consisting of beams of the Internet of Things device, and an action of the reinforcement learning algorithm is a subset of the beam set consisting of beams selected from the beams of the Internet of Things device. Spatial resource identification system for internet of things networks.
2. The spatial resource identification system for an Internet of Things network according to claim 1 , characterized in that the Internet of Things device uses the available spatial resource to perform device-to-device communication with other Internet of Things devices.
3. The spatial resource identification system for Internet of Things networks according to claim 1 , wherein the compensation of the reinforcement learning algorithm is defined as a probability of searching for spectrum resources relative to the energy consumed to identify the Internet of Things device.
4. The spatial resource identification system for Internet of Things networks according to claim 1 , characterized in that the wireless resource harvesting edge applies a c-greedy algorithm for the action of the reinforcement learning algorithm.
5. The spatial resource identification system for an Internet of Things network of claim 1, characterized in that the wireless resource harvesting edge collects the spectrum identification information from the Internet of Things device, executes the reinforcement learning algorithm to calculate compensation for optimal beam set determination, shares identification parameters derived based on previously collected spectrum identification information and previously derived identification parameters with the Internet of Things device, and retrains and searches for optimal identification parameters based on results reported from the Internet of Things device.
6. When the operation of the radio resource harvesting edge becomes impossible, The Internet of Things devices are grouped and a leader and a follower are determined within the group; The follower implements a beam selection strategy that can maximize the sensing probability with the minimum sensing power using the information about the optimal beam received from the leader, and transmits information about the sensing power threshold of the beam generated by the beam selection strategy to the leader; The leader generates a control variable corresponding to information on an optimal beam that can prevent overlapping between beams by utilizing the information received from the followers, and transmits the control variable to each follower; The followers and the leaders repeatedly exchange information on their respective strategies, The spatial resource identification system for an Internet of Things network as described in claim 1, wherein the reader terminates the iteration when a preset objective function calculated via the received information reaches an inflection point where it decreases and increases, and at the end of the iteration, determines the beam having the received information as the beam to be applied for communication, and turns off the remaining beams.
7. An identification stage in which each of the plurality of Internet of Things devices searches for a vacant spatial resource in the vicinity where communication is possible, and the wireless resource harvesting edge derives an optimal spatial resource where communication is possible through reinforcement learning; a reporting step in which the plurality of Internet of Things devices send the results of searching for available spatial resources to the wireless resource harvesting edge via a control channel, and the wireless resource harvesting edge receives the results of searching for available spatial resources and sends the optimal communication-enabled spatial resources searched for through the reinforcement learning of the previous identification step to the Internet of Things devices; A communication step in which the plurality of Internet of Things devices perform device-to-device communication through the optimal communication-enabled spatial resource, and the wireless resource harvesting edge updates the result of searching for the available spatial resource received in the reporting step, and starts searching for the optimal communication-enabled spatial resource using reinforcement learning; A spatial resource identification method for an Internet of Things network, comprising:
8. When the operation of the radio resource harvesting edge becomes impossible, The Internet of Things devices are grouped to determine a leader and followers in the group; The follower implements a beam selection strategy that can maximize the sensing probability with the minimum sensing power using the information on the optimal beam received from the leader, and transmits information on the sensing power threshold of the beam generated by the beam selection strategy to the leader; The leader generates a control variable corresponding to information on an optimal beam that can prevent overlapping between beams by utilizing the information received from the followers, and transmits the control variable to each follower; The follower and the leader repeatedly exchange information on their respective strategies; The method for identifying spatial resources for an Internet of Things network as described in claim 7, further comprising the step of: the reader terminating the iteration when a preset objective function calculated via the received information reaches an inflection point where it decreases and increases, determining the beam having the received information at the time of termination as the beam to be applied for communication, and turning off the remaining beams.
Citation Information
Patent Citations
Sensing data compression device, sensing data restoration device, sensing data transmission device, sensing data compression program, sensing data restoration program and sensing data transmission program
JP2021083036A
Control information for data transmission to NB (Narrowband) devices and co-located MBB (Mobile Broadband) devices using common communications resources
JP2021501544A