A Dynamic Spectrum Access Method Based on Attention Mechanism and Echo State Network

Through the Att-DESN framework, the problems of low efficiency, poor adaptability and insufficient robustness of spectrum resource allocation in 6G networks are solved, efficient and intelligent spectrum management is achieved, user collision rate is reduced, service quality is improved, and service quality is improved, and heterogeneous network environment is adapted.

CN120263321BActive Publication Date: 2025-08-01HANGZHOU DIANZI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510736531.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-08-01
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

The prior art has low efficiency in spectrum resource allocation, poor adaptability and insufficient robustness in 6G networks. Especially in high dynamic environments, the problem of overestimation of Q value is serious, and it cannot effectively solve the focus on selective spectrum characteristics, resulting in limited decision-making accuracy.

Method used

The dual echo state network (Att-DESN) framework based on attention enhancement is adopted, combining the additive attention mechanism and the echo state network (ESN), and by dynamically allocating the importance of spectrum characteristics, designing a mixed reward function and QoS scoring model, expanding the path loss model and introducing the target network soft synchronization mechanism to achieve efficient and intelligent spectrum management.

Benefits of technology

Significantly reduce the collision rate between secondary users and primary users, improve service quality scores, adapt to heterogeneous network environments, improve spectrum utilization and system throughput, improve training convergence speed and computing efficiency, and adapt to flexibility in different application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263321B_ABST
    Figure CN120263321B_ABST
Patent Text Reader

Abstract

The present invention discloses a dynamic spectrum access method based on the attention mechanism and the echo state network. First, a 6G space-ground integrated network system model and a three-dimensional network topology model are constructed, and the state space and the action space are defined. Secondly, the Bahdanau attention mechanism is introduced into the DDQN framework, and combined with the short-term memory characteristics of the ESN, a hybrid architecture with time series modeling ability is constructed. Then, a reward function that takes into account both spectrum efficiency and PU protection is designed, and the network QoS is evaluated through a weighted multi-index model. Finally, based on the Markov state space and the action space, the dynamic spectrum characteristics are processed by the hybrid architecture, combined with the reward function and the QoS evaluation mechanism, and an offline-online two-stage optimization strategy is adopted to complete the autonomous dynamic spectrum access of the space-ground integrated network. The present invention is compatible with the heterogeneous requirements of terrestrial networks, satellite networks, and the Internet of Things, and achieves significant performance improvement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - field of wireless communication technology and artificial intelligence technology, and particularly relates to a dynamic spectrum access (DSA) method based on attention - enhanced deep reinforcement learning, especially a dynamic spectrum allocation method based on an attention - enhanced deep Q - network (DQN). Background Art

[0002] The next - generation wireless network faces the challenges of scarce and fragmented spectrum resources. With the rapid growth of IoT devices and the research and deployment of 6G networks, the wireless communication network is facing an increasingly tense spectrum resource problem. It is predicted that the number of global connected devices will exceed 200 billion by 2030. The existing static spectrum allocation method can no longer meet the explosive growth of spectrum demand. As a key solution to solve the spectrum scarcity problem, DSA technology allows SUs to intelligently share licensed spectrum resources without interfering with PUs.

[0003] Traditional DSA methods are mainly divided into three categories: one is the distributed coordination method based on game theory, which has low decision - making efficiency and is difficult to converge; the second is the perception - based heuristic algorithm, which has poor adaptability to environmental dynamic changes; the third is the early reinforcement learning method, such as Q - learning, which has certain learning ability but faces the problem of state - space explosion. In recent years, deep reinforcement learning (DRL) has been introduced into the DSA field, such as DQN and its variants. Although it partially solves the problem of high - dimensional state processing, there are still the following technical bottlenecks: the over - estimation problem of Q - value leads to policy deviation; the decision - making reliability is insufficient in a partially observable environment; the adaptability to heterogeneous user scenarios is poor; the convergence speed is slow and the training stability is insufficient.

[0004] In addition, after the 6G network introduces millimeter - wave and terahertz frequency bands, the signal attenuation intensifies and the environmental interference is significant. There is an urgent need for a more intelligent spectrum allocation strategy. For example, when ultra - reliable low - latency communication coexists with enhanced mobile broadband, DSA needs to improve the system throughput while ensuring high reliability. The performance of existing methods in a partially observable environment still needs to be improved, especially in the highly dynamic scenario of satellite communication and ground network coexistence. Although the existing patent CN117896027A uses the DQN framework to achieve spectrum allocation, there are still two key defects in practical applications: firstly, this scheme uses the traditional DQN structure and fails to effectively solve the over - estimation problem of Q - value in a high - dynamic environment; secondly, its state representation completely depends on the original observation data and lacks a selective attention mechanism for spectrum features, resulting in limited decision - making accuracy in a partially observable scenario. Therefore, there is an urgent need for an innovative solution that can solve the above problems simultaneously. Summary of the Invention

[0005] The objective of the present invention is to provide a DSA method and system based on an attention-enhanced dual echo state network (Att-DESN) to address the problems of low spectrum allocation efficiency, poor adaptability, insufficient robustness, and limited future network compatibility in the prior art. By combining the attention mechanism with the double deep Q-network (DDQN) framework of the echo state network (ESN), the present invention solves the problems of scarce and fragmented spectrum resources in 6G and Internet of Things (IoT) networks and realizes efficient and intelligent spectrum management.

[0006] The present invention is applicable to future heterogeneous network environments, including scenarios where terrestrial communication, satellite communication, and a large number of IoT devices coexist, and can significantly reduce the collision rate between secondary users and primary users (SU-PU), improve the quality of service (QoS) score, and provide key technical support for dynamic spectrum sharing in 6G networks.

[0007] The present invention has achieved a technological breakthrough through the following core innovations:

[0008] (1) The additive attention mechanism is creatively combined with ESN. The importance of spectrum features is dynamically allocated through attention weights, and at the same time, the temporal processing ability of ESN is used to capture the dynamic characteristics of the spectrum.

[0009] (2) A hybrid reward function and a QoS scoring model are proposed to optimize the reliable and high-QoS communication objectives on the premise of ensuring the hard constraint of the collision rate of PUs.

[0010] (3) In view of the characteristics of the 6G space-ground integrated network, the path loss model is extended and a target network soft synchronization mechanism is introduced, which significantly improves the policy convergence and robustness under high-attenuation channels.

[0011] The technical solution adopted by the present invention is: a DSA method based on an attention-enhanced dual deep echo state network, including the following steps:

[0012] Step 1: Construct a 6G space-ground integrated network system model: establish a distributed DSA scenario including N primary users (PU) and M secondary users (SU), where each PU exclusively occupies a wireless channel, and the SU acts as an autonomous agent to dynamically adjust the access strategy by sensing the channel state; construct a three-dimensional network topology model, including a three-layer heterogeneous architecture of low-earth orbit satellite nodes, terrestrial base stations, and mobile terminals.

[0013] Step 2: Design the state space and action space: define the observation state space of the SU as a Markov process of the PU activity state, and the action space as a channel selection or waiting decision;

[0014] Step 3: Integrate the attention mechanism and ESN: Introduce the Bahdanau attention mechanism into the DDQN framework to dynamically focus on the key features of the input data, enhance the model's ability to process time-varying and noisy spectral environments; at the same time, combine the short-term memory characteristics of ESN to construct a hybrid architecture with temporal modeling capabilities;

[0015] Step 4: Optimize the reward function and QoS scoring: Design a reward function that takes into account both spectral efficiency and PU protection, and evaluate the network QoS through a weighted multi-index model;

[0016] Step 5: Training and deployment: Based on the Markov state space and discrete action space constructed in Step 2, process the dynamic spectral features through the attention-enhanced dual ESN architecture in Step 3, combine the multi-objective reward function and QoS evaluation mechanism designed in Step 4, and adopt an offline-online two-stage optimization strategy: In the offline stage, use experience replay and dual network synchronous training to achieve policy convergence, and in the online stage, deploy a lightweight model to achieve real-time spectrum decision-making, and finally complete the autonomous dynamic spectrum access in the space-ground integrated network environment.

[0017] Preferably, in Step 3, an attention-enhanced dual echo state network is used to achieve feature extraction and decision optimization of the spectrum state. The Bahdanau attention mechanism processes the input observations through a three-layer structure: First, the spectral features are represented as the hidden representation obtained through linear transformation and activation by the tanh function, where the learnable parameter is initialized using a normal distribution, where ensures the sparsity of feature selection; subsequently, the feature importance is calculated through the scoring function and the attention weights are generated using the softmax function; the enhanced state retains the original information through a residual connection, where represents element-wise multiplication. This enhanced state is input into the dual echo state network, and its dynamic equation is . The input weight matrix projects the attention-enhanced state to the reservoir layer, and the reservoir weight matrix is , and its spectral radius is set to 0.6 to ensure the echo state property. The output weight matrix is a trainable output layer for generating Q-value estimates. Under the structure of DDQN, action selection and evaluation are implemented through separate and networks, and the target value , where, represents the immediate reward, is the discount factor, is the target network, is the evaluation network, is the next state, is the next action, are the parameters of the evaluation network, are the parameters of the target network. Finally, the output of the double echo state network with attention enhancement is , where is the current state-action pair.

[0018] Preferably, in step 4, the reward function is designed in four cases according to the SU behavior: the reward when successfully accessing is proportional to the logarithm of the channel capacity: ; a penalty related to the interference level is imposed when there is a collision between SUs , the function uses the X logarithm form to map the interference value to the interval to ensure that the penalty intensity is proportional to the interference degree. The fixed penalty when there is a collision between the PU and the SU is ; the reward is 0 when no channel is selected. Among them, represents the secondary user at time selects the channel the signal-to-interference-plus-noise ratio (SINR), which is determined by the channel gain , the primary user interference power , the interference power between secondary users and the thermal noise power together, and the specific calculation formula is . Among them is the transmission power of the primary user, represents the path loss between the primary user and the secondary user, represents the channel state (1 for occupied, 0 for idle). is the channel bandwidth, is the power spectral density of the thermal noise.

[0019] Preferably, in step 4, the QoS score is achieved by weighted combination of multiple key performance indicators, and its expression is , where each component is normalized. represents the normalized average access success rate of the SU, and the calculation method is , where is the total number of access attempts, represents the th attempt is successful. and respectively represent the average collision rates after normalization for PU and SU, and the calculation methods are similar. The calculation formula for the PU collision rate is represents the collision event with PU. represents the average throughput after normalization, and its original value is determined by the channel bandwidth , signal-to-interference-plus-noise ratio , noise loss factor and the training batch duration jointly, and the calculation formula is , where is the SINR of SU on the channel . represents the normalized value of the long-term reward, and its original value is obtained by discounting and accumulating the immediate rewards, and the calculation formula is , where is the discount factor, is the immediate reward of SU at time , is the time range. The weight coefficient satisfies , and can be adjusted according to the specific application scenario. For example, in the ultra-reliable low-latency communication scenario, (the weight of the SU-PU collision rate) can be increased to preferentially guarantee the quality of service of PU.

[0020] Preferably, in step 5, the training process adopts the optimization strategy of deep reinforcement learning, and the specific implementation includes three key components: the experience replay buffer used to store the transition samples , and the sample correlation is broken by randomly sampling a small batch of data to improve the training efficiency; the exploration rate adopts a linear decay strategy, starting from the initial value and gradually decaying to . This design enables the agent to fully explore the environment in the initial stage of training and gradually tend to utilize the learned strategies in the later stage; the learning rate is set to during the training process, and the discount factor . These hyperparameters have been verified through a large number of experiments to obtain the best performance. The capacity of the experience replay buffer is set to , and each time training, samples are uniformly sampled from for network parameter update, and this mechanism effectively alleviates the training instability problem caused by the data sequence correlation.

[0021] Preferably, in step 5, the online deployment system implements the following functions: The real-time spectrum sensing module collects channel state observations at a fixed period (typical value is 10 ms). These observations include the occupancy status of each channel (0 indicates idle, 1 indicates occupied by the primary user), SINR estimation values and other key information; after receiving the observations, the attention mechanism module first calculates the hidden representation , where is a learnable parameter matrix, is the hidden layer dimension, is the -th feature component of the observation; subsequently, through attention weight calculation and normalization to obtain the enhanced state ; the decision engine inputs the enhanced state into the trained Att-DESN network, evaluates the value of each action through the Q function , and selects the optimal channel access action , where represents the network parameters; the system maintenance mechanism includes real-time monitoring of channel characteristic change indicators , when exceeds a preset threshold (typical value is 0.3), it automatically triggers the model fine-tuning process, and this threshold is determined according to historical data statistics to ensure that the system always adapts to the current wireless environment.

[0022] Preferably, in step 5, the training environment is configured as: Based on the python3.12 and Anaconda environments, using libraries such as tensorflow, numpy, and os for configuration.

[0023] The present invention has the following characteristics and beneficial effects:

[0024] High spectrum utilization and low collision rate: Through the attention-enhanced DDESN framework, the present invention achieves significant performance improvement in complex scenarios such as 6G networks and the Internet of Things.

[0025] Cross-network architecture compatibility: The unique channel modeling method is compatible with the heterogeneous requirements of terrestrial networks, satellite networks, and the Internet of Things. By integrating the WINNER II path loss model and the atmospheric attenuation term unique to satellite communication, the system can simultaneously handle spectrum management in millimeter wave (30 - 300 GHz), terahertz (0.1 - 10 THz), and traditional UHF bands. Atmospheric attenuation term, the system can simultaneously handle spectrum management in millimeter wave (30 - 300 GHz), terahertz (0.1 - 10 THz), and traditional UHF bands.

[0026] Fast convergence and computational efficiency: Combining the advantages of ESN and DDQN, the training convergence speed is increased by 57.1% compared to traditional Q-learning. The fixed random reservoir design of ESN reduces the computational complexity to Moreover, the sparsification process of the attention mechanism reduces the communication overhead by more than 30%, making it suitable for deployment on IoT devices with limited computing resources.

[0027] Scalable system architecture: The modularly designed reward function and QoS scoring system support flexible parameter adjustment. By modifying the weight coefficients it can be adapted to different application scenarios such as industrial Internet of Things (low latency priority) and smart cities (high reliability priority). The system throughput model supports bandwidth configurations from 36 MHz to 1 GHz, meeting the diverse requirements from narrowband Internet of Things to enhanced mobile broadband. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The technical solutions of the present invention will be described in detail below with reference to the drawings. It should be noted that the drawings are only used to exemplarily describe the embodiments of the present invention to facilitate those skilled in the art to better understand the technical solutions of the present invention. Without departing from the core idea of the present invention, those skilled in the art can obtain other implementation manners according to the drawings.

[0029] Figure 1 is the training flowchart of the DSA system according to the embodiment of the present invention;

[0030] Figure 2 is the schematic diagram of the space-air-ground integrated system model according to the embodiment of the present invention;

[0031] Figure 3 is the schematic diagram of the three-dimensional spatial distribution of PUs and SUs in the embodiment of the present invention.

[0032] The specific implementation details of the above drawings can be further understood in conjunction with the embodiment part of the present invention. Those skilled in the art can adjust the system configuration and parameters shown in the drawings according to the actual application scenarios, but these adjustments are all within the protection scope of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0033] In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0034] As Figure 1As shown in the figure, the training process of the Att-DESN model of the present invention includes steps such as environmental perception, attention enhancement, network training, and model update: First, the agent weights the historical observations through the attention mechanism to generate enhanced observations; then the observations are input into the echo state network for state expansion and stored in the experience replay pool together with actions and rewards; subsequently, data is sampled from the experience replay pool to update the DDQN-ESN network parameters; finally, the agent interacts with entities such as low-orbit satellites, PU / SU terminals, etc. in the Markov spectrum environment to obtain new observations and rewards, realizing the iterative optimization of the strategy.

[0035] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "transverse", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation of the present invention. In addition, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0036] In the description of the present invention, the term "attention mechanism" refers to the Bahdanau attention mechanism, which is used to dynamically weight input features; the term "spectrum sensing" refers to obtaining the channel state information of primary users and secondary users through wireless signal detection technology; the term "echo state network" refers to a recurrent neural network structure with fixed random weights, which is used to process time series data.

[0037] Embodiment 1

[0038] This embodiment provides a DSA method for future 6G networks, and its system architecture is as Figure 2 shown. Figure 2 It is a schematic diagram of the system model of the present invention, in which the low-orbit satellite senses the spectrum environment, the PU transmitter sends an ideal signal, the SU transmitter generates a potential interference signal, the PU terminal and the SU receiver receive and feedback the channel state, and the agent outputs actions to access available channels according to the sensing results and attention-enhanced representations, and obtains corresponding rewards according to the PU / SU conflict situation.

[0039] This method combines the Bahdanau attention mechanism with DDQN and ESN to achieve efficient spectrum resource allocation. The specific implementation steps are as follows:

[0040] Step 1, system initialization phase.

[0041] In the system initialization phase, network parameters are first set, including the reference distance , path loss exponent , channel gain coefficient and . At the same time, communication parameters are configured, the PU transmission power dBm, the SU transmission power dBm, the channel bandwidth is MHz, the attention mechanism parameters are initialized to follow the normal distribution, and a replay buffer D with a capacity of 300 pieces of experience data is established.

[0042] Step 2, spectrum sensing and state representation.

[0043] During the spectrum sensing and state representation process, the SU monitors the occupancy status of channels in real time through a wireless receiver. The system state space is constructed, where represents the occupancy status of the th channel. Considering a 20% sensing error probability , the environment is modeled as a partially observable Markov decision process.

[0044] Step 3, attention enhancement processing.

[0045] In the attention enhancement processing phase, the Bahdanau attention mechanism is applied to the observed state . First, the hidden state is calculated, then the attention score is calculated, and the attention weight is obtained through the softmax function. Finally, the enhanced state representation is output.

[0046] Step 4, decision-making and action execution.

[0047] When making a decision and executing an action, is input into the DDQN-ESN network, where the ESN part uses for state update. The action is selected through the -greedy strategy, and the optimal action is selected with probability , and an available action is randomly selected with probability . The channel access action is executed, where 0 means staying idle, Indicates which channel is selected.

[0048] Step 5: Reward calculation and learning phase.

[0049] In the reward calculation and learning phase, the immediate reward is calculated according to four situations : When successfully accessing, the reward is proportional to the logarithm of the channel capacity; when there is a conflict between SUs, a penalty is imposed according to the interference level; when there is a conflict between PU and SU, a fixed penalty -C3 is imposed; when no channel is selected, the reward is 0. Store the experience into the replay buffer D, and sample batch data from D to update the network parameters.

[0050] Step 6: Performance evaluation phase.

[0051] When evaluating the performance, calculate the QoS score, comprehensively considering the access success rate , the PU conflict rate , the SU conflict rate , the normalized throughput and the long-term cumulative reward and other indicators, and use the weight coefficient to perform weighted summation to obtain the final QoS score.

[0052] Figure 3 It is a schematic diagram of the spatial distribution of the PU transmitter, PU receiver, SU transmitter, and SU receiver in the X-Y-Z plane in the simulation scenario. This figure shows the random distribution of the transmitters (Transmitter) and receivers (Receiver) of PUs and SUs within a space of 50km × 50km × 500km. A total of 8 PUs and 4 SUs are arranged in the test scenario. Circles and triangles in the figure represent transmitters and receivers respectively, and the coordinates are in kilometers, which are used to construct the path loss model and evaluate the spectrum allocation performance at different positions.

[0053] Specifically, the technical effects of this embodiment are verified through simulation experiments: in the test scenario containing 6 PUs and 2 SUs, compared with the current advanced DEQN method, the present invention reduces the SU-PU conflict rate by 70% and improves the QoS score by 7.2%. In the more complex 8PU-4SU scenario, it still maintains the advantage of a 33% reduction in the conflict rate, and at the same time, the convergence speed is increased by 57.1% compared with the Q-learning method. These improvements are mainly due to the DDQN structure effectively alleviating the problem of overestimation of Q values, the attention mechanism enhancing the ability to extract key features, and the ESN architecture improving the efficiency of processing time-series data.

[0054] Table 1

[0055]

[0056] The proposed Att-DESN framework of the present invention can be further extended and applied to other wireless resource management scenarios, including but not limited to fields such as power control, user access selection, and network slice resource allocation. The Att-DESN training process is shown in Table 1. By adjusting the feature extraction method of the attention mechanism and the design of the reward function, this framework can adapt to the dynamic decision-making requirements in different communication scenarios, demonstrating good generality and scalability. In particular, in intelligent communication systems that need to handle partially observable environments, the method proposed in the present invention has significant technical advantages and application potential.

[0057] It should be understood that the above embodiments are only the preferred embodiments of the present invention. For those skilled in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

[0058] Embodiment 2

[0059] This embodiment provides a DSA system based on Att-DESN, including one or more processors and a storage module.

[0060] The storage module is used to store program codes and model parameters. When the program is executed by the processor, the processor can call the pre-trained Att-DESN model to make DSA decisions based on the real-time simulated wireless environment observation data, and update the model parameters to achieve adaptive optimization.

[0061] This embodiment also provides a computer-readable storage medium, which stores a computer program that can implement the above DSA method when executed by the processor. The storage medium includes media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0062] The above has described the embodiments of the present invention in detail in conjunction with the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principle and spirit of the present invention, various changes, modifications, substitutions, and variations to these embodiments including components still fall within the protection scope of the present invention.

Claims

1. A dynamic spectrum access method based on the attention mechanism and the echo state network, characterized in that It includes the following steps: Step 1, construct a 6G space-ground integrated network system model; Step 2, define the observation state space of the secondary user SU in the 6G space-ground integrated network system model as a Markov process of the primary user PU's activity state, and the action space as channel selection or waiting decision; Step 3, introduce the Bahdanau attention mechanism in the DDQN framework to dynamically focus on the key features of the input data, and at the same time combine the short-term memory characteristics of the ESN to construct a hybrid architecture with time series modeling ability; Step 4, design a reward function that takes into account both spectrum efficiency and PU protection, and evaluate the network QoS through a weighted multi-index model; Step 5, based on the Markov state space and action space, process the dynamic spectrum features through the hybrid architecture, combine the reward function and the QoS evaluation mechanism, and adopt an offline-online two-stage optimization strategy to complete the autonomous dynamic spectrum access in the space-ground integrated network environment.

2. The dynamic spectrum access method based on the attention mechanism and the echo state network according to claim 1, wherein The 6G space-ground integrated network system model includes a distributed DSA scenario with N primary users PU and M secondary users SU, where each PU exclusively occupies a wireless channel, and the SU, as an autonomous intelligent agent, dynamically adjusts the access strategy by sensing the channel state; The three-dimensional network topology model of the 6G space-ground integrated network system model includes a three-layer heterogeneous architecture of low-earth orbit satellite nodes, ground base stations, and mobile terminals.

3. The dynamic spectrum access method based on the attention mechanism and the echo state network according to claim 2, wherein The specific implementation process of Step 3 is as follows: The Bahdanau attention mechanism processes the input observations through a three-layer structure: First, the spectral features are linearly transformed and activated by the tanh function to obtain the hidden representation ; Subsequently, the feature importance is calculated through the scoring function , and the softmax function is used to generate the attention weights ; The enhanced state is the product of the spectral features and the corresponding attention weights, and the original information is retained through the residual connection; Enhance the state input double echo state network to obtain a hidden representation , which serves as the input data under the DDQN framework.

4. The dynamic spectrum access method based on the attention mechanism and the echo state network according to claim 3, characterized in that The reward function described in step 4 is designed in four cases according to the SU behavior: the reward at successful access is proportional to the logarithm of the channel capacity: , where represents the signal-to-interference-plus-noise ratio SINR when the secondary user k selects channel c at time t; a penalty related to the interference level is imposed during the collision between SUs , and the function adopts the X-logarithm form to map the interference value to the interval ; the fixed penalty during the collision between the PU and the SU is ; the reward is 0 when no channel is selected.

5. The dynamic spectrum access method based on the attention mechanism and the echo state network according to claim 4, characterized in that In step 4, the QoS is achieved by weighted combination of multiple performance metrics, and the expression is , where each component has been normalized; represents the average access success rate after normalization of the SU; and respectively represent the average collision rates after normalization of the PU and the SU; represents the average throughput after normalization; represents the normalized value of the reward, and the weight coefficient satisfies .

6. The dynamic spectrum access method based on the attention mechanism and the echo state network according to claim 5, wherein The specific implementation process of Step 5 is as follows: The training process adopts an optimization strategy of deep reinforcement learning, and an experience replay buffer is used to store transition samples , where represents the state, represents the action, represents the immediate reward, is the next state; the exploration rate adopts a linear decay strategy; the value of each action is evaluated through the Q function to select the optimal channel access action , where represents the network parameters.

Citation Information

Patent Citations

  • Distributed dynamic spectrum allocation method and device based on deep reinforcement learning

    CN117896027A

  • Dynamic spectrum and power control method based on cyclic Flash reinforcement learning

    CN118870549A