System and method for providing dynamic spectrum access
The system addresses scalability and adaptability challenges in LEO satellite networks by using edge computing and deep learning algorithms with blockchain, ensuring efficient and secure spectrum access for broadband services.
Patent Information
- Application Number
- PCT/SG2025/050482
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-17
- Filing Date
- 2025-07-16
- Publication Date
- 2026-01-22
AI Technical Summary
Existing dynamic spectrum access methods struggle to adapt to complex IoT environments in LEO satellite networks, facing challenges such as scalability, real-time adaptability, and computational bottlenecks, which hinder efficient broadband services.
A system combining edge computing, deep learning algorithms, and blockchain technology to optimize spectrum usage, utilizing graph neural networks and deep Q-networks for dynamic spectrum access, with a reputation-based consensus mechanism to enhance security and efficiency.
The system provides low-latency, reliable broadband services by efficiently utilizing idle spectrum resources, improving transmission security and enhancing spectrum usage efficiency in LEO satellite IoT networks.
Smart Images

Figure SG2025050482_22012026_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD FOR PROVIDING DYNAMIC SPECTRUM ACCESSCROSS-REFERENCE TO RELATED APPLICATION
[0001] The application claims the benefit of priority of Singapore patent application No. 102024021 13X, filed 17 July 2024, the content of it being hereby incorporated by reference in its entirety for all purposes.TECHNICAL FIELD
[0002] Various aspects of this disclosure relate to methods, server apparatuses, and systems for providing dynamic spectrum access, and in particular methods, server apparatuses, and systems for low earth orbit (LEO) satellite related applications.BACKGROUND
[0003] The following discussion of the background art is intended to facilitate an understanding of the present disclosure only. It should be appreciated that the discussion is not an acknowledgement or admission that any of the material referred to was published, known or is part of the common general knowledge of the person skilled in the art in any jurisdiction as of the priority date of the disclosure.
[0004] The rapid proliferation of Internet of Things (loT) devices has intensified demand for efficient spectrum utilization in wireless communications. Traditional dynamic spectrum access (DSA) methods, which rely on rule-based optimization algorithms or game-theoretic approaches, struggle to adapt to complex loT environments characterized by excessive device access, channel reductions, and dynamic channel -quality deterioration. For instance, spectrumbroker models for cellular networks and combined MAC-physical layer techniques for cognitive radio networks exhibit limitations in scalability and real-time adaptability.
[0005] Deep learning solutions, such as convolutional neural networks (CNN), have been adopted for primary-user detection. Long short-term memory (LSTM) networks have been utilized for spectrum prediction, improved detection accuracy but introduced challenges. Such challenges may include low sample efficiency, discrete action-space constraints, and computational bottlenecks when scaling to heterogeneous loT networks. While deep learning solutions such as graph neural networks (GNN) may be used to model node relationships and spectrum spatiotemporal features, they typically face performance degradation with large-scale graphs and global information processing. Similarly, deep Q-networks (DQN) enableintelligent spectrum decisions through reinforcement learning but may suffer from sample inefficiency and limited action-space exploration.
[0006] loT devices may be found in low earth orbit (LEO) satellite networks. To compete for LEO satellite spectrum and orbit resources, the construction of 6G communication networks based on LEO satellite communication infrastructure has triggered. How to achieve a larger commercial share in the future 6G network era is worthy of attention for industry around the world. The upstream industrial chain of LEO satellite industry is dominated by big powers, while the downstream ground application level has a larger market and industrial space. The LEO satellite system represented by Starlink has gradually begun to be networked in the Southeast Asian region since 2023. The supporting LEO satellite industrial applications are still in a state of scarcity, and the satellite commercial market is facing a reshuffle, which means a good opportunity for commercial companies to grab.
[0007] Unlike traditional GEO (geostationary) satellites, mobile satellite modules, and other models that focus on narrowband services and mainly support short data or positioning information, LEO satellite networks have opened the door to broadband services, paving the way for satellite services for broadband applications such as video and high-definition image transmissions. However, it is pertinent to provide an improved dynamic spectrum access for such broadband applications in order to provide users with reliable broadband services with acceptable quality.
[0008] There exists a need to provide an improved technical solution for dynamic spectrum access, particularly for, but not limited to, LEO satellite networks related applications.SUMMARY
[0009] The present disclosure was conceptualized to provide or build a low latency, highly reliable communication mechanism, and combine Al technology with edge computing to provide one or more users with high-quality broadband services.
[0010] In some embodiments, there may include a system oriented to terrestrial broadband loT services based on LEO satellites, supporting long range (LoRa) transmission protocol, compatible with 4G and 5G networks, independent power supply, Beidou satellite antenna and Wifi protocols, etc. In some embodiments, edge-computing-based image recognition, dynamic spectrum access and blockchain technologies may be used to reduce satellite bandwidth occupation, improve transmission security and enhance spectrum usage efficiency. Such anarrangement may be utilized to solve spectrum and tariff issues in the process of satellite broadband application.
[0011] In some embodiments, there is provided a dynamic spectrum access algorithm to better support broadband satellite loT services. The dynamic spectrum access algorithm may include a self-learning mechanism for dynamic spectrum access based on deep learning, fully utilizing the idle spectrum resources in the surrounding loT environment, solving the bandwidth resource congestion problem in the widespread broadband business application of future satellite loT, thus ensuring the broadband sendee quality of the system.
[0012] In some embodiments, there may comprise a LEO satellite loT blockchain solution for the enhancement of network fusion security. Facing the future situation of multiple ground loT integration, blockchain technology is adopted to ensure the security of data transmission for different types of terrestrial loTs.
[0013] In some embodiments, the blockchain includes a consensus algorithm to facilitate good execution and strike a balance between energy loss and system security. The consensus algorithm may be a reputation based consensus algorithm.
[0014] According to an aspect of the present disclosure there is provided a system for providing dynamic spectrum access, the system comprising: one or more sensor nodes configured to collect environmental data; a plurality of Internet of things (loT) gateways arranged in data or signal communication with the one or more sensor nodes via a network protocol, each loT gateway comprising at least one spectrum channel; a distributed ledger configured for secure data processing between each of the plurality of loT gateways and the one or more sensor nodes; and a dynamic spectrum access module arranged in data or signal communication with the plurality of loT gateways, the dynamic spectrum access module configured to implement a deep learning algorithm, the deep learning algorithm configured to obtain the environmental data as input, parse the environmental data to obtain spectrum environment parameters, predict one or more use states of the at least one spectrum channel of each of the plurality of loT gateways; and provide a spectrum access action based on the predicted one or more use states.
[0015] In some embodiments, the deep learning algorithm may comprise a graph neural network (GNN) component and / or a graph theory coloring network component, fused with a deep Q network (DQN) component, wherein the GNN component and / or the graph theory coloring network component is configured to predict the one or more use states of the at leastone spectrum channel, and the DQN component is configured to provide the spectrum access action based on a reinforcement learning algorithm.
[0016] In some embodiments, the DQN component defines an action space A = {0, 1, ..., N} where N represents a number of available channels, and wherein an agent selects at least one channel from the number of available channels to access for communication at each time step.
[0017] In some embodiments, the dynamic spectrum access module is configured to train the GNN component and / or the graph theory coloring network component and the DQN component using a replay memory technique.
[0018] In some embodiments, the dynamic spectrum access module is configured to represent or model the DQN component as a Bellman equation for reinforcement learning, the Bellman equation comprising a plurality of target Q-values forming a Q-table.
[0019] In some embodiments, each of the plurality of target Q-values forming the Q-table are updated using a temporal difference learning approach.
[0020] In some embodiments, each of the plurality of Internet of things (loT) gateways further comprise a solar-energy based power supply and a communication module.
[0021] In some embodiments, the system further comprises a Starlink terminal device, the Starlink terminal device configured to establish communication between the plurality of loT gateways and a LEO satellite.
[0022] In some embodiments, the network protocol comprises a long range (LoRa) protocol.
[0023] In some embodiments, the distributed ledger comprises a private blockchain network. The private blockchain network may be configured to implement a reputation-based consensus mechanism for any new loT gateway joining the system.
[0024] In some embodiments, the one or more use states of the spectrum channel comprises at least one of a connection successful state, a channel busy state, and a connection failed state.
[0025] In some embodiments, the one or more use states comprise a plurality of node feature vectors and adjacency matrices.
[0026] According to another aspect of the present disclosure there is provided a method for providing dynamic spectrum access for a LEO satellite network, the LEO satellite network comprising a plurality of loT gateways arranged in data or signal communication with one or more sensor nodes via a network protocol; comprising the steps of: collecting, from the one or more sensor nodes, environmental data; obtaining, by a dynamic spectrum access modulecomprising a deep learning algorithm, the environmental data; parsing, using the dynamic spectrum access module, the environmental data to obtain spectrum environment parameters; predicting, using the deep learning algorithm, one or more use states of at least one spectrum channel of each of the plurality of loT gateways; and providing a spectrum access action based on the predicted one or more use states; wherein the method comprises securing data processing between each of the plurality of loT gateways and the one or more sensor nodes using a distributed ledger.
[0027] In some embodiments, the deep learning algorithm comprises a graph neural network (GNN) component and / or a graph theory coloring network component, fused with a deep Q network (DQN) component, wherein the predicting the one or more use states of the spectrum channel is based on the GNN component and / or graph theory coloring network component, and the providing the spectrum access action is based on the DQN component and a reinforcement learning algorithm.
[0028] In some embodiments, the method further comprises defining, based on the DQN component, an action space A = {0, 1, ..., N] where N represents a number of available channels, and wherein an agent selects at least one channel from the number of available channels to access for communication at each time step.
[0029] In some embodiments, the method further comprises training the GNN component and / or the graph theory coloring network component, and the DQN component using a replay memory technique.
[0030] In some embodiments, the method further comprises representing or modelling the DQN component as a Bellman equation for reinforcement learning, the Bellman equation comprising a plurality of target Q-values forming a Q-table.
[0031] In some embodiments, the target Q-values forming the Q-table are updated using a temporal difference learning approach.
[0032] In some embodiments, the one or more use states of the spectrum channel comprises at least one of a connection successful state, a channel busy state, and a connection failed state.
[0033] In some embodiments, the one or more use states comprise a plurality of node feature vectors and adjacency matrices.
[0034] In some embodiments, the distributed ledger is optionally a blockchain network. The distributed ledger may be a private blockchain network. The private blockchain network may be configured to implement a reputation-based consensus mechanism for any new loT gateway joining the LEO satellite network.
[0035] According to another aspect of the present disclosure there is provided a non- transitory computer-readable medium storing computer executable code comprising instructions for providing dynamic spectrum access according to any one of the methods as described.
[0036] According to another aspect of the present disclosure there is provided a method for providing dynamic spectrum access for a LEO satellite network, the LEO satellite network comprising at least one loT gateway arranged in data or signal communication with one or more sensor nodes via a network protocol; comprising the steps of: collecting, from the one or more sensor nodes, environmental data; obtaining, by a dynamic spectrum access module comprising a deep learning algorithm, the environmental data; parsing, using the dynamic spectrum access module, the environmental data to obtain spectrum environment parameters; predicting, using the deep learning algorithm, one or more use states of at least one spectrum channel of each of the at least one loT gateway; and providing a spectrum access action based on the predicted one or more use states.
[0037] According to another aspect of the present disclosure there is provided a system for providing dynamic spectrum access, the system comprising: one or more sensor nodes configured to collect environmental data; at least one Internet of Things (loT) gateways arranged in data or signal communication with the one or more sensor nodes via a network protocol, each loT gateway comprising at least one spectrum channel; and a dynamic spectrum access module arranged in data or signal communication with the at least one loT gateway, the dynamic spectrum access module configured to implement a deep learning algorithm, the deep learning algorithm configured to obtain the environmental data as input, parse the environmental data to obtain spectrum environment parameters, predict one or more use states of the at least one spectrum channel of each of the at least one loT gateway; and provide a spectrum access action based on the predicted one or more use states.BRIEF DESCRIPTION OF THE DRAWINGS
[0038] The disclosure will be better understood with reference to the detailed description when considered in conjunction with the non-limiting examples and the accompanying drawings, in which:FIG. 1A is a schematic block diagram of a server apparatus for providing dynamic spectrum access.- FIG. IB is a schematic block diagram comprising the server apparatus of FIG. 1 A, with one or more sensor nodes and loT gateways, implemented as a system for providing dynamic spectrum access.- FIG. 2 is a schematic diagram illustrating a LEO satellite loT system incorporating the system of FIG. IB.- FIG. 3 illustrates an embodiment of the one or more sensor nodes and sensor flow information diagram.- FIG. 4 illustrates an embodiment of a dynamic spectrum access solution based on a hybrid DQN and GNN network.- FIG. 5 illustrates an embodiment of a private permissioned blockchain network configured to provide secure data processing between a plurality of loT gateways and the sensor nodes.- Fig. 6 illustrates a system model comprising a plurality of mobile loT devices applying for access channels to a central server, wherein the central server is configured to interact with a model layer to provide spectrum access action (e.g. channel access strategy) based on a dynamic access method.- FIG. 7 illustrates a learning framework of the model layer of FIG. 6.- FIG. 8 is a table showing simulation parameter settings of maximum network benefits, in the form of a Collaborative MaxSum Reward (CMSR) criterion or the protocol-based criterion for maximum impoverished cognitive user network benefits coloring, in the form of a Collaborative Max Proportional Fair (CMPF) criterion, graph coloring scheme model, and random coloring scheme model, according to some embodiments of the present disclosure.- FIG. 9 is a graph illustrating network benefits comparison between DQN-CMSR and CMSR in 100 iterations.- FIG. 10 is a graph illustrating network benefits comparison between DQN-CMPF and CMPF in 100 iterations.- FIG. 11 is a graph illustrating network benefits comparison between DQN-CMSR and CMSR in 200 iterations.- FIG. 12 is a graph illustrating network benefits comparison between DQN-CMPF and CMPF in 200 iterations.- FIG. 13 is a graph illustrating network benefits comparison between DQN-CMSR and CMSR with changing coverage.- FIG. 14 is a graph illustrating network benefits comparison between DQN-CMPF and CMPF with changing coverage.- FIG. 15 is a graph illustrating network benefits comparison between DQN-CMSR and CMSR with changing channels.- FIG. 16 is a graph illustrating network benefits comparison between DQN-CMPF and CMPF with changing channels.- FIG. 17 is a graph illustrating network benefits comparison between DQN-CMSR and CMSR with changing secondary users.- FIG. 18 is a graph illustrating network benefits comparison between DQN-CMPF and CMPF with changing secondary users.- FIG. 19 is a graph illustrating network benefits comparison between DQN-CMSR and CMSR with changing primary users.- FIG. 20 is a graph illustrating network benefits comparison between DQN-CMPF and CMPF with changing primary users.- FIG. 21 is a flowchart showing a generalized method for providing dynamic spectrum access, according to some embodiments of the present disclosure.DETAILED DESCRIPTION
[0039] The following detailed description refers to the accompanying drawings that show, by way of illustration, specific details, and embodiments in which the disclosure may be practiced. These embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosure. Other embodiments may be utilized, and structural and logical changes may be made without departing from the scope of the disclosure. The various embodiments are not necessarily mutually exclusive, as some embodiments can be combined with one or more other embodiments to form new embodiments.
[0040] Features that are described in the context of an embodiment may correspondingly be applicable to the same or similar features in the other embodiments. Features that are described in the context of an embodiment may correspondingly be applicable to the other embodiments, even if not explicitly described in these other embodiments. Furthermore, additions and / orcombinations and / or alternatives as described for a feature in the context of an embodiment may correspondingly be applicable to the same or similar feature in the other embodiments.
[0041] In the context of various embodiments, the articles “a”, “an” and “the” as used with regard to a feature or element include a reference to one or more of the features or elements. As used herein, the term “and / or” includes any and all combinations of one or more of the associated listed items.
[0042] While such terms as "first," "second," etc., may be used to describe various elements, such elements must not be limited to the above terms. The above terms are used only to distinguish one element from another, and do not define corresponding elements, for example, an order and / or significance of the elements. Without departing from the scope of rights of the specification, a first element may be referred to as a second element, and similarly, the second element may be referred to as the first element.
[0043] As used herein, the term “data” may be understood to include information in any suitable analog or digital form, for example, provided as a file, a portion of a file, a set of files, a signal or stream, a portion of a signal or stream, a set of signals or streams, and the like. The term data, however, is not limited to the aforementioned examples and may take various forms and represent any information as understood in the art.
[0044] As used herein, the term “processor” refers to a circuit, including analog circuits, digital circuits, or hybrid circuits, or their constituent components. Any other kind of implementation of the respective functions which will be described in more detail below may also be understood as a “circuit” in accordance with an alternative embodiment. A digital circuit may be understood as any kind of a logic implementing entity, which may be special purpose circuitry or a processor executing software stored in a memory, or a firmware.
[0045] As used herein, the term “module” refers to, forms part of, or includes an Application Specific Integrated Circuit (ASIC); an electronic circuit; a combinational logic circuit; a field programmable gate array (FPGA); a processor (shared, dedicated, or group) that executes code; other suitable hardware components that provide the described functionality; or a combination of some or all of the above, such as in a system-on-chip. The term module may include memory (shared, dedicated, or group) that stores code executed by the processor. A single module or a combination of modules may be regarded as a device. A processor may include one or more modules. For example, multiple modules described in this disclosure may form a processor.
[0046] As used herein, the term “associate”, “associated”, and “associating” indicate a defined relationship (or cross-reference) between two items.
[0047] As used herein, “memory” may be understood as a non-transitory computer-readable medium in which data or information can be stored for retrieval. References to “memory” included herein may thus be understood as referring to volatile or non-volatile memory, including random access memory (“RAM”), read-only memory (“ROM”), flash memory, solid-state storage, magnetic tape, hard disk drive, optical drive, etc., or any combination thereof. Furthermore, it is appreciated that registers, shift registers, processor registers, data buffers, etc., are also embraced herein by the term memory. It is appreciated that a single component referred to as “memory” or “a memory” may be composed of more than one different type of memory, and thus may refer to a collective component including one or more types of memory. It is readily understood that any single memory component may be separated into multiple collectively equivalent memory components, and vice versa. Furthermore, while memory may be depicted as separate from one or more other components (such as in the drawings), it is understood that memory may be integrated within another component, such as on a common integrated chip.
[0048] As used herein, the term “device” may be understood to refer to any apparatus, equipment, or component, whether standalone or integrated, that performs a specific function or set of functions. This includes, but is not limited to, mechanical, electrical, electronic, optical, or electromechanical systems, subsystems, and assemblies. A device may comprise one or more components, modules, or units that are designed to interact with each other to achieve a particular purpose.
[0049] As used herein, the term “low earth orbit (LEO) satellite” may refer to an artificial satellite that orbits the Earth at an altitude within a predetermined distance, for example 2,000 kilometers of the Earth's surface, i.e., in relatively lower Earth orbit, which is much closer to the Earth compared to satellites in medium or high Earth orbits. In some embodiments, LEO satellites may be used to enable relatively high-speed, low-latency broadband Internet and data transmission, making them useful for connecting remote or underserved locations, supporting mobile communications (such as satellite phones), and providing reliable connectivity for business operations, events, and emergency response centers. LEO satellites may utilize one or more spectrum channels (e.g. frequency bands) for communication, depending on their application and technical requirements. The selection of spectrum channel may depend on operational requirements, regulatory constraints, and desired trade-offs between bandwidth, coverage, equipment size, and susceptibility to atmospheric effects.
[0050] As used herein, the term “dynamic spectrum access” may be understood to refer to a process of wireless communication where devices dynamically access underutilized portions of a spectrum, such as a radio-frequency spectrum, enabling efficient shared use. This approach differs from traditional fixed spectrum allocation by allowing opportunistic or coordinated access to spectrum resources in real-time based on availability of one or more spectrum channels. In LEO satellite applications, dynamic spectrum access (DSA) may refer to adaptive methods for allocating underutilized radio-frequency spectrum to satellite and terrestrial users, addressing the unique challenges of high mobility, doppler shifts, and interference in low Earth orbit networks. Non-limiting examples of DSA may include real-time spectrum sensing, machine learning optimization, and federated coordination.
[0051] As used herein, the term “Internet of things (IoT)” broadly refers to a network of interconnected devices — such as sensors, meters, vehicles, machinery, and other equipment — that collect, transmit, and exchange data using a communication channel, for example LEO satellite links, as the primary or supplementary communication channel. In some embodiments, IoT in the context of LEO satellites may include the deployment and operation of distributed, sensing, and control devices that utilize LEO satellite constellations to achieve data connectivity for one or more applications. In some embodiments, the IoT devices may send and receive real-time data via LEO satellites, supporting time-sensitive applications in various applications. In some embodiments, support for low-power wide-area networks (LPWAN) may be achieved using wireless communication protocols such as LoRa and Long Range-Frequency Hopping Spread Spectrum (LR-FHSS).
[0052] As used herein, the term “distributed ledger” broadly refers to a digital system for recording, sharing, and / or synchronizing data, where the ledger may be replicated and maintained across multiple sites, nodes, or institutions, rather than being stored in a single, centralized location. The distributed ledger may comprise multiple nodes in a network, each node in the network holds an identical copy of the ledger and participates in validating and updating records through consensus algorithms, ensuring that all copies remain consistent and tamper-resistant. Non-limiting examples of a distributed ledger may include blockchains, directed acyclic graphs (DAGs), Hyperledger fabric, etc.
[0053] As used herein, the term “sensor node” includes any device, apparatus, system, and / or software component that detects, measures, monitors, or records physical, environmental, or operational conditions, phenomena, or properties, and generates output indicative of those conditions. The output may be in the form of electrical, mechanical, optical,or other signals, and may be processed by hardware or software systems for further analysis or control purposes. A sensor node may include both hardware components (e.g., transducers, detectors, circuits) and software components (e.g., algorithms, data processing modules) that together enable the detection, measurement, and interpretation of the desired parameters. Some non-limiting examples of sensors include image capturing sensors, for example, cameras, and light detection and ranging sensor (lidar).F0054] As used herein, the term “configured to” broadly refers to the design, arrangement, or adaptation of a system, device, component, or module to perform a specific function or achieve a particular outcome. The term includes both hardware and software implementations wherein in a hardware implementation, the physical components are arranged, programmed, or structured to carry out the intended function(s), and in the context of programming and software, a device is operable under executable instructions (e.g., software, firmware) to perform the specified function(s) when executed by one or more processors. The resultant configuration allows the system or component to perform the stated function, either inherently or after suitable programming or activation, without requiring substantial modifications to its structure or operational logic.
[0055] According to various embodiments, a circuit may include analog circuits or components, digital circuits or components, or hybrid circuits or components. Any other kind of implementation of the respective functions which will be described in more detail below may also be understood as a "circuit" in accordance with an alternative embodiment. A digital circuit may be understood as any kind of a logic implementing entity, which may be special purpose circuitry or a processor executing software stored in a memory, firmware, or any combination thereof. Thus, in various embodiments, a "circuit" may be a digital circuit, e.g., a hard-wired logic circuit or a programmable logic circuit such as a programmable processor, e.g., a microprocessor (e.g., a Complex Instruction Set Computer (CISC) processor or a Reduced Instruction Set Computer (RISC) processor). A "circuit" may also include a processor executing software, e.g., any kind of computer program, e.g., a computer program using a virtual machine code such as e.g., Java.
[0056] According to an aspect of the disclosure and with reference to FIG. 1A, there is provided a server apparatus for providing dynamic spectrum access. The server apparatus may be suitable for the provision of dynamic spectrum access to spectrum channels of an LEO satellite system. In some embodiments, the server apparatus may be configured to implement a deep learning algorithm, the deep learning algorithm configured to obtain environmental dataas input, parse the environmental data to obtain spectrum environment parameters, predict one or more use states of the spectrum channels; and provide a spectrum access action based on a prediction of one or more use states.
[0057] The server apparatus may comprise a processor and a memory, the processor is capable of being configured to execute instructions stored in the memory to provide dynamic spectrum access. In the embodiment illustrated in FIG. 1A, the server apparatus may be a communications server apparatus. The communications server apparatus may be in the form of a server computer 10, the server computer 10 may be a single server as illustrated schematically in FIG. 1A, or have the functionality performed distributed across multiple server components.
[0058] In some embodiments, the server computer 10 includes a communication interface 12 (e.g. configured to receive data, such as data obtained from one or more sensor nodes). The communication interface 12 may include a transmitter module and / or a receiver module allowing the server computer 10 to communicate over a communications network. The communication interface 12 may include one or more user-interfaces configured to provide users for user control and may include, for example, one or more computing peripheral devices such as display monitors, computer keyboards and the like.
[0059] The server computer 10 may further include a processor in the form of processing unit 14 and a memory 16. The memory 16 may be used by the processing unit 14 to store, for example, data to be processed, including environmental data obtain from one or more sensor nodes, and any associated data, metadata, etc.
[0060] As shown in FIG. IB, there is a system 100 for providing dynamic spectrum access, utilizing the server computer 10. In operation, the processing unit 14 of the server computer 10 may be configured to receive, for example, from one or more computer devices 15, a request for dynamic spectrum access 1001, for one or more spectrum channels of at least one loT gateways associated with the LEO satellite system. The system 100 may comprise one or more sensor nodes 17, each sensor node 17 configured to collect environmental data 1002, one or more Internet of Things (loT) gateways 18, the loT gateways 18 arranged in data or signal communication with the one or more sensor nodes 17 via a network protocol, each loT gateway 18 comprising at least one spectrum channel. In some embodiments, the environmental data 1002 may be processed to a suitable format 1002A, or may be sent directly, via a network. The system 100 may further comprise a distributed ledger 21, the distributed ledger 21 configured for secure data processing between each of the plurality of loT gateways 18 and the one ormore sensor nodes 17. In some embodiments, the distributed ledger may include a blockchain network, the blockchain network implementing a consensus algorithm 1005 to facilitate good execution and strike a balance between energy loss and system security.
[0061] The processing unit 14 may be, or form part of a dynamic spectrum access module, and may be arranged in data or signal communication with the plurality of loT gateways 18. The dynamic spectrum access module may be configured to implement a deep learning algorithm, the deep learning algorithm configured to obtain the environmental data as input, parse the environmental data to obtain spectrum environment parameters, predict one or more use states of the at least one spectrum channel of each of the plurality of loT gateways; and provide a spectrum access action 1004 based on the predicted one or more use states 1003.
[0062] In some embodiments, the deep learning algorithm comprises a graph neural network (GNN) component and / or a graph theory coloring network component, fused with a deep Q network (DQN) component, wherein the GNN component and / or the graph theory coloring network component is configured to predict the one or more use states of the at least one spectrum channel, and the DQN component is configured to provide the spectrum access action based on a reinforcement learning algorithm.
[0063] As illustrated in FIG. IB, the data flow may be facilitated by a network 180. The network 180 may be an internal network, such as an Intranet, or may be the Internet. In some embodiments, the network 180 may form part of a cloud network.
[0064] In some embodiments, the system of FIG. IB may be utilized or adopted in a LEO satellite loT system, as illustrated in the system 200.
[0065] As shown in FIG. 2, the system 200 comprises a plurality of sensor nodes 202, a plurality of loT gateways 204, a Starlink terminal device 206, a LEO satellite 208, a cloud server 210, user terminals 212, such as mobile phone(s) and computer terminal(s).
[0066] In some embodiments, data storage and communication between each of the plurality of loT gateways 204 may be secured using a distributed ledger, such as a blockchain. The plurality of loT gateways 204 may be configured to support one or more communication module.
[0067] The Starlink terminal device 206 may be configured to establish communication between each of the plurality of loT gateways 204 and the LEO satellite 208.
[0068] The network protocol used to facilitate communication between the plurality of sensor nodes 202 and the plurality of loT gateways 204 may comprise a long range (LoRa) protocol to facilitate a relatively larger transmission range. In some embodiments, thedistributed ledger may comprise a private blockchain network or private permissioned blockchain network. In some embodiments, the one or more use states of the spectrum channel may comprise at least one of a connection successful state, a channel busy state, and / or a connection failed state. The one or more use states may be represented, or comprises, a plurality of node feature vectors and adjacency matrices.
[0069] FIG. 3 shows a system 300 for hardware and sensor flow of information or data between the loT sensor nodes 202, a satellite gateway 204, a satellite dish 310, a database 320, and front-end interface 330. The loT sensor nodes 202 may be configured to obtain environmental parameters or data 3001, such as, but not limited to, air flow related data (for e.g. wind), humidity related data (for e.g. relative humidity), and temperature related data.
[0070] The obtained data 3001 may be sent to the satellite gateway 204 via a LoRa-enabled device, and in turn be transmitted to the satellite dish 310. The database 320 may be arranged in data communication with the satellite dish 310 to facilitate automated exchange, storage, and / or processing of data received from or sent to satellites.
[0071] The front-end interface 330 may be configured to obtain real-time data from one or more users, and the obtained real-time data may be received by the satellite dish 310. In some embodiments, the real-time data may include one or more requests for historical data. The obtained historical data (e.g. from the database 320) may be transmitted from the satellite dish 310 for display on the front-end interface 330.
[0072] In some embodiments, the LoRa protocol supports interconnection with 4G or 5G cellular networks, WiFi modules, deep-learning-based image recognition, solar independent power supply, mobile applications, and web-based monitoring programs. In some embodiments, the LoRa protocol may be configured to simultaneously support single hop relay forwarding transmission, which may be used when sensor nodes are positioned relatively far apart.
[0073] In some embodiments, the plurality of loT gateways 204 may include 4G-based, 4G compatible, 5G-based, and / or 5G compatible communication modules, wherein the different environmental data can also be transmitted back through cellular networks within a geographical location, such as in a city. The various communication modules may be utilized to maintain compatibility with existing mainstream cellular networks.
[0074] In some embodiments, one or more of the plurality of loT (satellite) gateways 204 may be configured to enable wireless data connectivity between various loT devices and LEO satellite ground stations by integrating one or more WiFi modules compliant with wirelessstandards, such as IEEE 802.11 standards. Such an arrangement allows, for example, the environmental data from the loT sensor nodes 202 and user data collected by the gateway to be transmitted over a local wireless network to the gateway, which then relays the data to the LEO satellite ground station for uplink to the satellite constellation. The use of WiFi modules facilitates flexible deployment of loT sensor nodes and end-user terminals within the coverage area of the gateway, reducing wiring complexity and supporting rapid installation in diverse environments. This architecture enables seamless aggregation of data from multiple distributed devices, which is then efficiently transmitted via the LEO satellite network for global connectivity, which may be particularly valuable in remote or infrastructure -limited regions.
[0075] In some embodiments, the loT gateways 204 may be embedded with graphics processing units capable to support artificial intelligence applications, such as, but not limited to, NVIDIA processing units. The graphics processing unit may enable on-device execution of deep learning models for image recognition and analysis. In some embodiments, the hardware acceleration may support real-time processing of visual data from connected image processing devices such as cameras or sensors, allowing for advanced functionalities such as object detection, classification, and anomaly recognition at the edge. By performing image processing locally, the system may reduce the bandwidth required for transmitting raw image data over the satellite link, thereby optimizing network utilization and enabling immediate decisionmaking based on recognized events.
[0076] In some embodiments, to facilitate autonomous operation and adaptability to remote or off-grid environments, the system 200 may be equipped with photovoltaic solar panels and an integrated power management subsystem. Such a configuration may provide a renewable and independent energy source, enabling continuous operation of the gateway, sensors, and communication modules without reliance on external electrical infrastructure. The solar- powered design may enhance system reliability and sustainability, making it particularly suitable for deployment in challenging loT scenarios such as environmental monitoring, agriculture, and disaster response.
[0077] In some embodiments, dedicated monitoring software, accessible via both mobile applications and web-based interfaces, may be developed to provide users with real-time visibility into the operational status of the plurality of sensor nodes 202 and gateways 204. The programs may allow remote users to monitor sensor readings, device health, and network connectivity, and to receive alerts or notifications regarding system events. The dual-platformapproach ensures accessibility for a wide range of users and devices, enhancing system manageability and user engagement.
[0078] FIG. 4 shows an embodiment of a system 400 for providing dynamic spectrum access solution based on a hybrid DQN and GNN network.
[0079] The system 400 may include a spectrum environment module 402, which provides states and rewards 4001 based on actions taken by the system 400. The spectrum environment may represent an available wireless communication spectrum, which may be subject to dynamic variations due to competing users, interference, and regulatory constraints.
[0080] The spectrum environment module 402 may take one or more actions based on a current dynamic spectrum access strategy or policy 404, which are fed back into the spectrum environment module 402.
[0081] The system 400 includes a memory buffer 406, the memory buffer 406 configured to store one or more use states, actions, rewards, and next state tuples.
[0082] The memory buffer 406 may be configured to interact with a mini-batch component 408, the mini-batch component 408 configured to receive the use states and rewards data from the memory buffer 406 as input, and pre-process the input for training to a dynamic spectrum access module 412. The dynamic spectrum access module 412 may comprise a fused Deep Q- Network (DQN) module and a Graph Neural Network (GNN) module. The fused DQN module and GNN module 414 may be trained or updated using a loss function 416, the loss function 416 may relate to the self-learning mechanism of the dynamic spectrum access module 412. As illustrated, the updated parameters output from the loss function may be propagated from a previous iteration 414A of the fused DQN module and GNN module to a current iteration 414B. The output of the fused DQN module and GNN module 416 may be a spectrum access action, which may include one or more dynamic spectrum access strategy.
[0083] The DQN modules may be configured to process the input states and generate Q- values for action selection, while the GNN modules may be configured to extract relational or topological features from the state inputs, such as the interactions or dependencies among multiple agents or spectrum channels.
[0084] In some embodiments, the joint model may be configured to receive data from the environment, processes these through the DQN and GNN, and outputs a dynamic spectrum access strategy.
[0085] It may be appreciated that the memory buffer 406 stores states, actions, rewards, and next state tuples data. This allows the system 400 to employ a replay memory technique, which enhances training stability and data efficiency.
[0086] From the memory buffer, mini-batches of experiences are sampled and passed into the joint model for training. A loss function is computed by comparing predicted Q-values with target values derived from the reward feedback and updated Q-values. The loss function-based parameter updating may be based on reinforcement learning, configured to guide updates to both neural network components to ensure convergence towards optimal spectrum access policies.
[0087] It may be appreciable that reinforcement learning may stabilized and made more efficient by using stored experience tuples for batch training. It is contemplated that the system may operate in a closed feedback loop where actions influence environmental states, which are in turn used to improve subsequent decision-making.
[0088] In some embodiments, the dynamic spectrum access strategy may be continuously refined through the iterative loop of action selection, reward feedback, and model updating. This strategy aims to optimize spectrum utilization while avoiding interference, maintaining quality of service (QoS), and adhering to regulatory requirements.
[0089] In some embodiments, the system 400 jointly employs deep reinforcement learning and graph-based feature extraction to enhance decision-making in complex spectrum environments.
[0090] The dynamic spectrum access (DSA) mechanism, leveraging graph neural network algorithms, may be configured to optimize spectrum utilization among multiple gateways and loT devices. The software module may be utilized to dynamically allocate communication channels based on real-time network conditions, interference patterns, and device density, thereby improving spectral efficiency and minimizing communication conflicts. In some embodiments, loT gateways 204 may be added to the system, with performance testing of the DSA mechanism done for the increased number of gateways for validation of scalability and robustness under dense deployment conditions.
[0091] In some embodiments, to ensure secure, decentralized, and tamper-resistant interconnection of loT devices over the LEO satellite 208, the blockchain algorithm tailored for satellite-based loT communication may facilitate the recording and verification of device transactions and data exchanges, enhancing trust and data integrity across distributed gateways. Full-scale verification and performance evaluation of the blockchain solution may beconducted as the number of deployed loT gateways 204 increases, to enable assessment of throughput, latency, and consensus efficiency in a satellite-enabled loT environment.
[0092] In some embodiments, when the LEO satellite loT system works in unlicensed spectrum, the proposed dynamic spectrum access method may facilitate better use of the bandwidth resource to support the broadband applications.
[0093] It may be appreciable that the fusion algorithm of graph neural network (GNN) and deep Q network (DQN) may improve spectrum access efficiency while maintaining a good access success accuracy. It may be appreciable that compared with traditional DQN, the computation time of the fused DQN / GNN module 414 can be reduced by over 35%. In summary, the proposed fused algorithm first uses GNN to interact with the environment and predict one or more use states of the loT spectrum environment. Subsequently, automatic learning and optimization of spectrum access policies may be achieved by selecting the mobile loT user’s actions based on these predicted states using the DQN’s target network, experience playback, and reinforcement learning techniques.
[0094] FIG. 5 illustrates a private permissioned blockchain network that may be utilized to provide secured data transmission and storage between the various components or modules in the dynamic spectrum access system. In particular, when different loT gateways 204 need to be connected, the proposed blockchain algorithm can secure the connections. If only one loT with one gateway and multiple sensors designed by same producer is utilized, the blockchain solution will not be essential.
[0095] In some embodiments, a new blockchain method with a totally distributed framework and new consensus algorithm may be developed. Unlike traditional blockchain solutions, no fixed fusion center will be required to be set in order to conduct or implement the blockchain algorithm. In some embodiments, loT terminals with computing power may take turns to serve as edge nodes. The system model of the proposed blockchain method is as shown in FIG. 5.
[0096] One or more blockchain nodes 510 in the network structure serves as a platform for collaborative spectrum sensing decisions and allocation. Each of the plurality of loT gateways 502, operating independently under similar conditions, may form a blockchain network cluster 500. To ensure the security of newly added gateways 502 and maintain uniformity among the loT gateways 502 in the network cluster 500, a unified authentication standard may be established. Under the blockchain configuration, each gateway 502 that is joining the networkcluster 500 may be required to pass a pre-defined authentication rule before being admitted. These authentication rules can be defined autonomously by the builders of the blockchain.
[0097] In some embodiments, each gateway 502 joining the network may be required to use a consensus-based training algorithm provided by the blockchain and a commonly recognized central aggregation algorithm. When the model training is completed, gateways 502 must undergo authentication to gain permission to join the network. Gateways 502 that pass the authentication can participate in the next round of local training. This ensures that newly added gateways 502 do not disrupt the training of existing gateways in the network, while also allowing new gateways to immediately engage in distributed learning and obtain high-quality parameters during the aggregation phase. Each gateway joining the blockchain network cluster 500 must adhere to the blockchain's privacy protocol, interact with each other through the blockchain platform, and share their local model parameters. A leader gateway may be selected by the consensus mechanism to aggregate the neural network model parameters, quickly forming a global model for the network cluster. The leader gateway then distributes the weight parameters of the global model to all gateways in the blockchain for a next round of edge learning and intelligent DSA.
[0098] In some embodiments, a reputation-based consensus mechanism based on the scenario may be utilized. Each gateway 502 that is a member of a private permissioned blockchain network has a reputation token record. The reputation token record of each gateway may change for each round of dynamic spectrum access by terrestrial users. If there are no malicious behaviours of the gateway, the number of reputation tokens for that gateway is the original number of reputation tokens plus the newly acquired reputation tokens. Instead, if there is a malicious user in this access round, the number of reputation tokens for the gateway is the original number of reputation tokens minus the penalty incurred for the malicious behaviours. The reputation record of a gateway may be mathematically expressed in Equation (1), as follows.where Reps tdenotes the reputation token of the sth gateway at time t, VLaccdenotes the newly acquired reputation token based on the situation that no malicious behaviors appear and V™aldenotes the newly lost reputation token based on the situation that malicious behaviors appear.
[0099] In some embodiments, when the LEO satellite loT works in unlicensed spectrum, the proposed dynamic spectrum access method assists the proposed dynamic spectrum accesssystem better use of the bandwidth resource to support the broadband applications. When various kinds of loTs are connected, the proposed blockchain algorithm can secure the connections. If only one loT with one gateway and multiple sensors designed by same producer is utilized, the blockchain solution will not be essential.
[0100] In some embodiments, the algorithms used in dynamic spectrum access and blockchains are different. For example, deep learning algorithms are used in the dynamic spectrum access algorithms. Then, a protocol-based consensus algorithm may be used in the blockchain.
[0101] In some embodiments, the original DSA algorithm with the joint graph and DQN method is proposed. Compared with traditional methods, the proposed algorithm can better consider the interactions between various dynamic channels and improve the flexibility and adaptability of spectrum allocation in loT. The spectrum access accuracy can be improved when compared to traditional graph-theory-based solutions.
[0102] In some embodiments, in order to overcome the limitations of a single deep learning or graph theory approach, a dynamic spectrum access (DSA) method that integrates graph theory coloring models with deep Q networks (DQN) may be adopted. Firstly, an initial graph- theoretic model may be introduced to define feasible spectrum allocation by calculating an interference-free allocation matrix and a network benefits vector. A strategy pool is then created to filter inefficient strategies using labels and colors based on a Collaborative MaxSum Reward (CMSR) and Collaborative Max Proportional Fair (CMPF) criteria. The spectrum environment for the DQN is defined based on this strategy pool, and the DQN model is trained to optimize spectrum access strategies through the Q-table update algorithm based on a Bellman equation. Such an approach leverages the self-learning capability of DQN and the theoretical strengths of graph theory, enabling efficient spectrum utilization in complex environments. The interactions between channels and the distribution of spectrum resources may be better considered, thus improving the flexibility and adaptability of spectrum allocation. It is contemplated that the above integrated graph theory coloring method with DQN may better adapt to the dynamic changes of the spectrum environment, so as to avoid falling into the dilemma of local optimal solutions. Experimental results (shown in FIG. 9 to FIG. 20) show that, compared with the traditional graph theory and random coloring schemes, the integrated graph theory coloring method with DQN has significantly improved spectrum utilization and user experiences.
[0103] FIG. 6 shows a system model of the proposed integrated graph theory coloring model and DQN that are combined to optimize dynamic spectrum access in loT applications. Various mobile loT devices 602 apply for access channels to the server, which classifies the application information into primary and secondary user information 604 and stores the relevant parameter information in one or more databases, which may be SQL databases 606, then in turn upload the relevant parameter information to one or more edge servers 608 and finally transmits to one or more central servers 610. The one or more central servers 610 may comprise cloud servers and storage servers. The one or more central servers 610 may interact with a model layer 612, the model layer 612 implementing the proposed integrated graph theory coloring model and DQN, to return the best channel access strategy, and passes the corresponding strategy layer by layer to ensure that users can access the spectrum channel correctly and quickly.
[0104] The interaction between the one or more central servers 610 and the model layer 612 culminates in the learning framework for the optimal channel access strategy, as shown in FIG. 7.
[0105] In some embodiments, the process which is as illustrated in FIG. 7, the one or more central servers 610 passes basic parameters 702 to the graph-theoretic coloring model, which includes key information such as the number of primary users, the number of secondary users, the maximum communication coverage, the number of channels and the radius of the protection zone. Based on the received parameters, the graph-theoretic coloring model generates a portion of channel assignment strategies using either the protocol-based criterion for maximum network benefits (CMSR) or the protocol-based criterion for maximum impoverished cognitive user network benefits (CMPF) and stores these strategies in a strategy pool. Then, a portion of strategies are extracted from the strategy pool, and these strategies are reconstructed into the class of environments that can be handled by the DQN model, including the action space, the state space, and the corresponding rewards. The DQN model learns to train continuously by interacting with the environment and searching for optimal strategies in the whole state space. The resulting optimal channel access strategy is passed back to the central server for channel allocation.
[0106] According to the learning framework, the relevant processes may be divided into four learning phases: modeling graph theoretic coloring models 710, building strategy pools 720, modeling DQN models 730, and agent training 740. These processes are described in detail as follows.
[0107] In the modelling graph theory learning phase, based on the actual layout of the cognitive network, a graph-theoretic model 712 is constructed, which may consist of an available spectrum matrix, a network benefits matrix, an interference matrix, and an noninterference allocation matrix. The available spectrum matrix and the network benefits matrix can be derived from the relative positions of primary and secondary users. The spectrum matrix L represents the availability of each channel for each secondary user. The matrix element ln.m indicates whether channel m is available to the secondary user n. If ln,m= 1, it means channel m is available for user n, otherwise ln,m = 0, mathematically expressed in Equation (2), as follows.
[0108] A network benefits matrix B may be defined, and is mathematically represented in Equation (3) as follows. The matrix B represents the network benefits of each secondary user on each channel. The matrix elementindicates the network benefits of the secondary user n on channel m. The network benefits can be calculated based on factors such as channel quality, signal interference, and transmission rate.
[0109] An interference matrix C may be defined, and is mathematically represented in Equation (4) as follows. The matrix C describes the interference relationships between users. The matrix element cn,km indicates whether secondary users n and k will interfere with each other if they simultaneously use channel m. If cn,k,m = 1, it means users n and k will interfere on channel m, otherwise
[0110] A non-interference allocation matrix A may be defined, and is mathematically represented in Equation (5) as follows. The matrix A may be a binary matrix representing the allocation of channels. The matrix element an.m indicates whether channel m is allocated to secondary user n. The non-interference allocation matrix must satisfy two conditions: firstly, no channel should be simultaneously allocated to users that would interfere with each other; secondary users can only use available channels.
[0111] In some embodiments, each secondary user’s obtained network benefits is definedand the network benefits vector for all secondary users is formed and mathematically represented in Equation (6), as follows.
[0112] The set of all feasible spectrum allocation methods may be denoted as A(L, Cfvx.M- In some embodiments, the next step is to find the spectrum allocation method that maximizes a particular network benefits function from among all possible methods.
[0113] In some embodiments, the define coloring guidelines 714 of the graph theory model 712 may be based on choosing the network benefits functions as Max-Sum-Reward (MSR) and Max-Proportional-Fair (MPF) functions. The network benefits function based on MSR, denoted as UMSR(R\ may be mathematically represented in Equation (7), as shown below.
[0114] The MSR network benefits function aims to maximize the total network benefits across all secondary users. It is calculated by summing the network benefits r„ for each user n, where rnis the sum of the product of the allocation matrix an,mand the network benefits matrix hn,macross all channels m. Such an approach focuses on maximizing the overall network performance without considering fairness among the individual users.
[0115] The network benefits function based on MPF, denoted as UMPF(R), as follows.
[0116] The MPF network benefits function aims to achieve a balance between maximizing total network benefits and ensuring fairness among users. It does this by taking the logarithm of the network benefits rnfor each user n, where rnis calculated similarly to the MSR approach. The logarithmic function helps to balance the network benefits among users, preventing any single user from dominating the resource allocation. In making optimal allocations, the choice of vertex labeling rules needs to be based on the desired allocation objective. The size of the label reflects the value of the vertex determined by the allocation goal and the benefits weights, and the larger the value of the vertex, the higher the label. Each label corresponds to a color, and the criterion of the coloring algorithm is to prioritize the coloring of the most valuable vertices. In order to achieve different desired goals for different network benefits functions, different labeling methods will be used to satisfy the needs of network benefits functions. In this paper, two coloring guidelines are defined based on two network benefits functions, MSR and MPF, which are CMSR and CMPF guidelines. The CMSR coloring criterion is shown below.where Dn,mrepresents the number of secondary users causing interference to secondary user n on channel m, and Inrepresents the set of available channels for secondary user n. The CMSR criterion considers the overall system’s maximum network benefits and the trade-off with neighboring user conflicts. Due to its consideration of neighborhood impact, this criterion is collaborative.
[0117] The CMPF coloring criterion is mathematically expressed in Equation (8), where an,mrepresents the spectrum allocation completed prior to this allocation for secondary user n. The label is based on the spectrum allocation that was already completed before this allocation.The label is derived from the previously completed spectrum allocation, ensuring proportional fairness by considering the historical allocation. The color assignment (channel selection) aims to maximize the benefit-to-interference ratio for the current allocation.
[0118] In the building strategy pools in block 720, to construct the strategy pool, an initial filtering 722 of the generated channel access strategies 716 is required to eliminate the less effective strategies. Specifically, this may involve calculating the average network benefits of all strategies 716, filtering out strategies 716 with network benefits less than the average network benefits, and placing the remaining strategies into a strategy pool 724. Let the network benefits of the z-th strategy be Ut, and there are N strategies in total. The average network benefits U of all strategies is calculated as shown below, in Equation (12) as follows.( |2 )
[0119] After filtering out strategies with network benefits less than U, the set of remaining strategies P, mathematically expressed in Equation (13), is as follows.
[0120] All strategies in the set P are then placed into the strategy pool 724 for further use.
[0121] In the modelling DQN model block 730, after constructing the strategy pool 724, a subset of strategies 732, also resulting in a mini batch strategies, may be sampled to build an environment suitable for DQN training. The reward for each strategy is calculated based on the network benefits function.
[0122] A subset of strategies 732 may be randomly sampled from the strategy pool P (724).The subset may be denoted as Psampied, mathematically expressed in Equation (14) as follows.
[0123] In defining the environment 734, the environment 734 for DQN is constructed using the sampled strategies 732. Each state in the environment may correspond to a specific strategy Pi G ^sampled- An action space A may be defined. The action space A consists of all possible actions Atthat can be taken from each state St. Here, Pi denotes the specific strategy associated with a particular state St.
[0124] In calculating a reward calculation, the reward r, for each strategy Pi is calculated using the network benefits function UPI(R). The network benefits function is determined by the specific objective of the strategy, mathematically expressed in Equation (15), as follows.( |5 )
[0125] Once the environment is constructed with the sample strategies and the rewards are defined, the DQN training process can begin. The training process involves updating the O'values based on the rewards and the expected future rewards. The Q-values are updated using the following, mathematically expressed as Equation (16), as follows.(16)where Q(St,At) is the Q-value for taking action Atin state S,, r, is the immediate reward received after taking action At, y is the discount factor, Q'(5f+i,Aniax(5 / , (o')) is the maximum Q-value for the next state S / +i.
[0126] In some embodiments, the Q-table is updated using a Bellman equation, such as a combination of the immediate reward and the discounted future reward. The updated equation is mathematically expressed as Equation (17), as follows.where 6 is the learning rate, Q(ShAt) is the updated Q-value for taking action At in state St, Q'(JSt,At) is the previous Q-value for taking action Atin state St, rtis the immediate reward received, y is the discount factor, Q(5 / +i,Amax(56 cu')) is the maximum Q-value for the next state £+i.
[0127] In the agent training block 740, the DQN model is trained to derive the optimal channel allocation strategy by iteratively updating the Q-values using the above equations. The training process can be summarized as follows:1. Initialize Q-Table: Initialize the Q-table with arbitrary values.2. Experience Replay: Store experiences (S,, Ahrt, St+i) in a replay memory.3. Mini-Batch Sampling: Sample a mini-batch 738 of experiences from the replay memory 736.4. Q-value Calculation: For each experience in the mini-batch, calculate the Q-values using the immediate Q-value update equation.5. Q-Table Update: Update the Q-values in the Q-table using the Q-table update equation.6. Strategy Update: Use the updated Q-table to derive the optimal strategy.The DQN model then passes the corresponding channel allocation strategy to the graph theory coloring model to achieve the best results.
[0128] Algorithm block 742 — The system modeling algorithm is used to optimize DSA by combining graph theory coloring algorithms and DQN. The inputs to the algorithm include parameters such as the number of channels, minimum transmission rate, maximum communication coverage, number of secondary users, number of primary users, and protected area radius. The output is the optimized channel allocation strategy. In the first step, an initial graph-theoretic model is constructed. This may involve computing the non-interference allocation matrix A, which ensures that no channel is allocated to interfering users. Next, we calculate the utility for each secondary user using the formula ' » ~~where an,mrepresents the allocation of channel m to user n, and bn,m represents the network benefits or utility that user n gets from using channel m. The set of all feasible spectrum allocation methods A(L, C)N%M is then defined, which considers all possible ways to allocate channels to users, given the constraints on the spectrum and the interference matrix.Next, the coloring guidelines based on network benefits functions aimed at maximizing the system’s overall network benefits (MSR) and minimizing interference (MPF) are defined. For each secondary user n, the CMSR T’CMSROO and CMPF ( / CMPF(«) are calculated. Based on these calculated network benefits, labels and colors to vertices (users) are assigned. In another step, an initial set of channel access strategies are generated.The average network benefits1“ 3rS-t ‘ ■' across all strategies are calculated. Strategies with network benefitsi f are retained, while less effective strategies are filtered out. The remaining strategies form the strategy pool P.
[0129] The DQN environment may then be modeled by randomly sampling a subset of strategies / Sampled from the strategy pool P. The state space and action space for the DQN environment are defined based on the strategies that were sampled. For each strategythe reward r, = UPI (R) may be calculated using the network benefits functions defined earlier.
[0130] The training of the DQN model includes initializing the Q-table with arbitrary values and iteratively updating the Q-values based on sampled experiences. During each training iteration, we store the experience (St, at, rt, St+i) in the replay memory, sample a mini-batch from the replay memory, and use the immediate Q-value update equation and the Q-table update equation to update the Q-values and the Q-table, repeating this process until convergence is reached. Finally, we use the updated Q-table to derive the optimal channel access strategy. This optimal strategy is then applied to the graph coloring model to obtain the best channel, as shown in Algorithm 1.
[0131] Simulation experiments were carried out based on the model / system of FIG. 7 and Algorithm 1, and results summarized in FIG. 9 to FIG. 20. Detailed experiments were conducted to investigate the relationship between different network benefits with the number of model training, the maximum communication coverage, the number of available channels, the number of secondary users, and the number of primary users. Also, the method of adding random labels (denoted as RAND) by randomly generating labels distributed between [0, 1] at each point and randomly selecting a color from the color list for coloring is compared. In order to enhance the illustrative nature of the experiments, after drawing on the experimental methodology in References [1] and [2], three sets of model schemes were created for comparison: the graph-coloring scheme model, the model of the proposed algorithm 1, and the random-coloring scheme model. The metric used in the experiments is the average network benefits, which reflects the overall communication quality, resource utilization efficiency, and fairness in service distribution within the cognitive radio network. High network benefits encompasses aspects such as higher data transmission rates, reduced error rates, improved quality of service, and enhanced system performance including throughput, capacity, and latency optimization.
[0132] FIG. 8 shows the setting of simulation parameters. For better simulation and comparison experiments, the same parameter ranges to control the variables may be set based on CMSR and CMPF coloring criterion, graph coloring scheme model, and random coloring scheme model. The simulation parameters are set as follows: Different parameters, such as the number of secondary users, the number of primary users, the maximum communication coverage, and the number of channels, are considered. Under each set of parameters, large number of simulation experiments are carried out to ensure the reliability and stability of the results.
[0133] Results and Discussion — After a series of simulations, corresponding simulation diagrams were generated. FIG. 9 and FIG. 10 present the comparison of different spectrum allocation schemes under the CMSR and CMPF guidelines, respectively, over 100 iterations.
[0134] In FIG. 9, the schemes compared are DQN-CMSR (Deep Q-Network with CMSR guidelines), CMSR (using traditional graph theory approach), and RAND (random allocation). The DQN-CMSR scheme shows a stable and high network benefits index, significantly outperforming both the traditional CMSR and random allocation schemes. Similarly, FIG. 10 compares DQN-CMPF, CMPF, and RAND schemes. The DQN-CMPF scheme maintains a consistently high network benefits index, surpassing both the CMPF and random allocationschemes. These figures highlight the superiority of DQN-based schemes in achieving higher utility under both guidelines over a shorter iteration period.
[0135] FIG. 11 and FIG. 12 extend the comparison of spectrum allocation schemes under the CMSR and CMPF guidelines to 200 iterations. FIG. 11 shows the results for DQN-CMSR, CMSR, and RAND schemes, with DQN-CMSR continuing to exhibit superior performance with a stable and high network benefits index throughout the extended iterations. FIG. 12 displays the comparison for DQN-CMPF, CMPF, and RAND schemes. The DQN-CMPF scheme consistently achieves a higher network benefits index compared to the traditional CMPF and random allocation schemes, maintaining its advantage over the longer iteration period. These figures reinforce the effectiveness and robustness of DQN-based schemes over extended iterations.
[0136] The comparison of spectrum allocation schemes under CMSR and CMPF guidelines over different iteration periods demonstrates the clear superiority of DQN-based schemes. Both short-term (100 iterations) and long-term (200 iterations) experiments show that DQN-CMSR and DQN-CMPF achieve higher network benefits indexes with greater stability compared to traditional graph theory-based and random allocation methods. These results highlight the potential of combining DQN with graph theory models to enhance spectrum allocation efficiency and adaptability in loT environments. Further experiments will explore the maximum communication coverage variations, channel number variations, and the number of primary and secondary user variations, providing a comprehensive assessment of the proposed method’ s performance under different conditions.
[0137] The experimental results of the models built based on the two coloring criterion schemes, CMSR and CMPF, are compared with those of the graph theory and random coloring schemes by varying the maximum communication coverage.
[0138] FIG. 13 shows the comparison results under the CMSR criterion, and it can be seen that the DQN-CMSR scheme significantly outperforms both the CMSR and RAND schemes for different values of Dmax, and the network benefits increase with increasing Dmax. FIG. 14 illustrates the comparison results under the CMPF criterion, and again, the DQN-CMPF scheme outperforms the CMPF and RAND schemes for most values of Dmax.
[0139] Next, the simulation is set up to change the number of channels to achieve the comparison results. FIG. 15 shows the comparison results under the CMSR criterion, where the DQN-CMSR scheme significantly outperforms the CMSR and RAND schemes in terms of network benefits under different channel numbers, especially when the number of channelsincreases. FIG. 16 shows the comparison results under the CMPF criterion, where the DQN- CMPF scheme outperforms the CMPF and RAND when the number of channels increases. This suggests that the scheme combined with a deep Q-network can effectively optimize the spectrum access strategy and improve the network performance under different channel numbers.
[0140] FIG. 17 and FIG. 18 show the comparison when the number of secondary users changes. In FIG. 17, as the number of secondary users increases, the network gains of the DQN- CMSR scheme are always higher than those of the CMSR and RAND schemes, and the reduction is slower, indicating that it maintains better performance when the number of secondary users increases. In FIG. 18, the DQN-CMPF scheme significantly outperforms the CMPF and RAND schemes when the number of secondary users is small, but the network benefits gradually decrease as the number of secondary users increases. Nevertheless, the DQN-CMPF scheme still outperforms the other two schemes overall.
[0141] FIG. 19 and FIG. 20, from another aspect, show the comparison when the number of primary users varies. In FIG. 19, the DQN-CMSR scheme exhibits some fluctuations in network revenue when the number of primary users increases but still outperforms the CMSR and RAND schemes overall. In FIG. 20, the DQN-CMPF scheme exhibits a more stable trend in network gains when the number of primary users increases and outperforms the CMPF and RAND schemes in most cases. These results show that the schemes combining DQN are all effective in optimizing the spectrum access strategy and improving the network performance when coping with different changes in the number of primary users.
[0142] It may be appreciable that through a series of simulation experiments, various schemes under different conditions are compared, including changes in training iterations, maximum communication coverage, number of channels, and number of secondary and primary users. The proposed scheme consistently outperforms the conventional approach, achieving higher network benefits with fewer iterations and maintaining this advantage as iterations increase. It demonstrates superior adaptability in optimizing spectrum utilization, yielding greater network benefits in multi-channel environments. When the number of secondary users increases, its network benefits decline more slowly, showcasing excellent management and optimization capabilities. Additionally, it exhibits better stability and performance as the number of primary users changes. Overall, the proposed scheme significantly enhances optimization capability and adaptability, effectively improving spectrum utilization efficiency and offering a flexible solution to the DSA problem.FOO 143] According to another aspect of the present disclosure there is provided a method 900 for providing dynamic spectrum access for a LEO satellite network, the LEO satellite network comprising a plurality of loT gateways arranged in data or signal communication with one or more sensor nodes via a network protocol; comprising the steps of:
[0144] Step S902: collecting, from the one or more sensor nodes, environmental data;
[0145] Step S904: obtaining, by a dynamic spectrum access module comprising a deep learning algorithm, the environmental data;
[0146] Step S906: parsing, using the dynamic spectrum access module, the environmental data to obtain spectrum environment parameters;
[0147] Step S908: predicting, using the deep learning algorithm, one or more use states of at least one spectrum channel of each of the plurality of loT gateways; and
[0148] Step S910: providing a spectrum access action based on the predicted one or more use states.
[0149] Step S912: wherein the method comprises securing data processing between each of the plurality of loT gateways and the one or more sensor nodes using a distributed ledger.
[0150] In some embodiments, the deep learning algorithm comprises a fused graph neural network (GNN) component and / or a graph theory coloring network component, and a deep Q network (DQN) component, wherein the predicting the one or more use states of the spectrum channel is based on the GNN component and / or the graph theory coloring network component, and the providing the spectrum access action is based on the DQN component and a reinforcement learning algorithm.
[0151] In some embodiments, the method further comprises defining, based on the DQN component, an action space A = [0, 1, ..., N] where N represents a number of available channels, and wherein an agent selects at least one channel from the number of available channels to access for communication at each time step.
[0152] In some embodiments, the method further comprises training the GNN component and / or the graph theory coloring network component, and the DQN component using a replay memory technique.
[0153] In some embodiments, the method further comprises representing or modelling the DQN component as a Bellman equation for reinforcement learning, the Bellman equation comprising a plurality of target Q-values forming a Q-table.
[0154] In some embodiments, the target Q-values within the Q-table are updated using a temporal difference learning approach.
[0155] In some embodiments, the one or more use states of the spectrum channel comprises at least one of a connection successful state, a channel busy state, and a connection failed state.
[0156] In some embodiments, the one or more use states comprise a plurality of node feature vectors and adjacency matrices.
[0157] In some embodiments, the distributed ledger is optionally a blockchain.
[0158] It may be appreciable that where there comprise only one loT gateway, the distributed ledger may be optional.
[0159] According to another aspect of the present disclosure there is provided a non- transitory computer-readable medium storing computer executable code comprising instructions for providing dynamic spectrum access according to any one of the aforementioned method.
[0160] In the present disclosure, a spectrum allocation optimization scheme that combines DQN and graph theory is proposed. The combination or fusion leverages the self-learning capability of DQN and the theoretical strengths of graph theory, enabling efficient spectrum utilization in complex environments. After filtering out the poorly performing data generated by graph theory, a strategy pool is built, where some data may be sampled from the pool and construct it as environmental variables that can interact with the DQN model (such as state intervals, action intervals, reward values, etc.). Finally, the DQN model may be interacted to achieve the optimal channel access strategy. Simulation experiments compare our scheme with traditional graph theory and random color matching schemes, and the results show that the proposed scheme exhibits superior adaptability and optimization, consistently outperforming traditional methods across different scenarios.
[0161] In summary, the proposed DSA scheme integrating DQN’s self-learning with graph theory, significantly improves spectrum utilization efficiency and user experience, offering an effective solution to the DSA problem. The proposed disclosure may provide methods for optimizing future spectrum allocation strategies with substantial theoretical and practical significance.
[0162] By employing a graph coloring model, initial channel allocation strategies may be effectively generated, aiming at reduction of interference and improvement of spectrum efficiency. Then, a strategy filtering step may be designed to evaluate and retain the most promising initial strategies, forming a strategy pool.
[0163] In particular, the DQN is adopted to train and optimize the solution within the strategy pool, adapting to dynamic spectrum environments and enhancing overall performances.
[0164] The combination of graph theory and DQN ensures high-quality initial strategies and an efficient optimization process. The proposed method dynamically adjusts strategies based on environmental changes, thereby improving spectrum utilization.
[0165] It may be appreciable that the commercial prospects of LEO satellite loT systems are broad, including but not limited to agriculture, forestry, environmental protection, ocean, power grid, and border defense. The present DSA model / module (in combination with LEO satellite loT systems) addresses one or more of the following drawbacks.
[0166] (i) Traditional loT monitoring relies on cellular networks around cities, making it impossible to transmit data back in more remote areas.
[0167] (ii) Traditional GEO satellite loTs are limited by the large latency and high fees caused by the long transmission path, and are mainly used for narrowband services such as short messages and monitoring status information. LEO satellite networks have enabled broadband applications for ground loT, supporting high-definition image and video data transmission in various scenarios, and the application scope is greatly expanded compared to traditional modes.
[0168] (iii) LEO satellite networks have lower tariff, which can truly provide low-rail, high- quality, and low-cost broadband application services for customers in various industries.
[0169] (iv) The LEO satellite networks represented by Starlink were connected to Southeast Asia for only around one year, and the related supporting industries are still in a blank state. Building a mature and reliable LEO satellite loT system can provide efficient solutions for environmental and commercial monitoring applications in Southeast Asia and the whole world.
[0170] The present disclosure of LEO satellite loT system based on dynamic spectrum access and blockchain improves the spectrum efficiency and security of the connections for various kinds of loTs.
[0171] (v) A dynamic spectrum access algorithm in loT has been proposed using a combination of joint graph theory and deep Q-Networks (DQN. The present disclosure adopts graph theory and DQN to enhance the channel access rate with lower computing power.
[0172] While the disclosure has been particularly shown and described with reference to specific embodiments, it should be understood by those skilled in the art that various changes in form and detail may be made therein without departing from the spirit and scope of thedisclosure as defined by the appended claims. The scope of the disclosure is thus indicated by the appended claims and all changes which come within the meaning and range of equivalency of the claims are therefore intended to be embraced.REFERENCES[11. F. Li, B. Shen, J. Guo, K.-Y. Lam, G. Wei, and L. Wang, Dynamic spectrum access for intemet-of-things based on federated deep reinforcement learning, IEEE Transactions on Vehicular Technology, vol. 71, Art. no. 7, 2022.[2], F. Li, L. Yang, J. Zhang, and L. Wang, A novel spectrum allocation scheme in femtocell networks using improved graph theory, in Springer, 2018, pp. 172-180.
Claims
1. CLAIMS1. A system for providing dynamic spectrum access, the system comprising: one or more sensor nodes configured to collect environmental data; a plurality of Internet of things (loT) gateways arranged in data or signal communication with the one or more sensor nodes via a network protocol, each loT gateway comprising at least one spectrum channel; a distributed ledger configured for secure data processing between each of the plurality of loT gateways and the one or more sensor nodes; and a dynamic spectrum access module arranged in data or signal communication with the plurality of loT gateways, the dynamic spectrum access module configured to implement a deep learning algorithm, the deep learning algorithm configured to obtain the environmental data as input, parse the environmental data to obtain spectrum environment parameters, predict one or more use states of the at least one spectrum channel of each of the plurality of loT gateways; and provide a spectrum access action based on the predicted one or more use states.
2. The system according to claim 1, wherein the deep learning algorithm comprises a graph neural network (GNN) component and / or a graph theory coloring network component, fused with a deep Q network (DQN) component, wherein the GNN component and / or the graph theory coloring network component is configured to predict the one or more use states of the at least one spectrum channel, and the DQN component is configured to provide the spectrum access action based on a reinforcement learning algorithm.
3. The system according to claim 2, wherein the DQN component defines an action space A = {0, 1, ..., N} where N represents a number of available channels, and wherein an agent selects at least one channel from the number of available channels to access for communication at each time step.
4. The system according to claim 2 or claim 3, wherein the dynamic spectrum access module is configured to train the GNN component and / or the graph theory coloring network component and the DQN component using a replay memory technique.
5. The system according to claim 4, wherein the dynamic spectrum access module is configured to represent or model the DQN component as a Bellman equation for reinforcement learning, the Bellman equation comprising a plurality of target Q-values forming a Q-table.
6. The system according to claim 5, wherein the target Q-values within the Q-table are updated using a temporal difference learning approach.
7. The system according to any one of claim 1 to claim 6, wherein each of the plurality of Internet of things (loT) gateways further comprise: a solar-energy based power supply and a communication module.
8. The system according to any one of claim 1 to claim 7, further comprising a Starlink terminal device, the Starlink terminal device configured to establish communication between the plurality of loT gateways and a LEO satellite.
9. The system according to any one of claim 1 to claim 8, wherein the network protocol comprises a long range (LoRa) protocol.
10. The system according to any one of claim 1 to claim 9, wherein the distributed ledger comprises a private blockchain network, and wherein the private blockchain network is configured to implement a reputation-based consensus mechanism for any new loT gateway joining the system.
11. The system according to any one of claim 1 to claim 10, wherein the one or more use states of the spectrum channel comprises at least one of a connection successful state, a channel busy state, and a connection failed state.
12. The system according to claim 11, wherein the one or more use states comprise a plurality of node feature vectors and adjacency matrices.
13. A method for providing dynamic spectrum access for a LEO satellite network, the LEO satellite network comprising a plurality of loT gateways arranged in data or signalcommunication with one or more sensor nodes via a network protocol; comprising the steps of: collecting, from the one or more sensor nodes, environmental data; obtaining, by a dynamic spectrum access module comprising a deep learning algorithm, the environmental data; parsing, using the dynamic spectrum access module, the environmental data to obtain spectrum environment parameters; predicting, using the deep learning algorithm, one or more use states of at least one spectrum channel of each of the plurality of loT gateways; and providing a spectrum access action based on the predicted one or more use states; wherein the method comprises securing data processing between each of the plurality of loT gateways and the one or more sensor nodes using a distributed ledger.
14. The method according to claim 13, wherein the deep learning algorithm comprises a graph neural network (GNN) component and / or a graph theory coloring network component, fused with a deep Q network (DQN) component, wherein the predicting the one or more use states of the spectrum channel is based on the GNN component and / or the graph theory coloring network component, and the providing the spectrum access action is based on the DQN component and a reinforcement learning algorithm.
15. The method according to claim 14, further comprising defining, based on the DQN component, an action space A = {0, 1, ..., N} where N represents a number of available channels, and wherein an agent selects at least one channel from the number of available channels to access for communication at each time step.
16. The method according to claim 14 or claim 15, further comprising training the GNN component and / or the graph theory coloring network component, and the DQN component using a replay memory technique.
17. The method according to claim 16, wherein further comprising representing or modelling the DQN component as a Bellman equation for reinforcement learning, the Bellman equation comprising a plurality of target Q-values forming a Q-table.
18. The method according to claim 17, wherein the target Q- values within the Q-table are updated using a temporal difference learning approach.
19. The method according to any one of claim 13 to claim 18, wherein the one or more use states of the spectrum channel comprises at least one of a connection successful state, a channel busy state, and a connection failed state.
20. The method according to claim 19, wherein the one or more use states comprise a plurality of node feature vectors and adjacency matrices.
21. The method according to any one of claim 13 to claim 20, wherein the distributed ledger is a private blockchain network, and wherein the private blockchain network is configured to implement a reputation-based consensus mechanism for any new loT gateway joining the LEO satellite network.
22. A non-transitory computer-readable medium storing computer executable code comprising instructions for providing dynamic spectrum access according to any one of the method of claims 13 to 21.
Citation Information
Patent Citations
Spectrum resource allocation method and system based on cooperative distributed DQN joint simulated annealing algorithm
CN113613332A
Dynamic spectrum access method of Internet of Things based on joint learning of graph neural network and deep Q network
CN118300724A
Dynamic spectrum access scheme based on federated learning, graph neural network and deep Q network
CN118741534A
Reinforcement learning for multi-access traffic management
US20220014963A1
Cited By
Industrial data acquisition method and system based on edge computing
CN121691192A