Dynamic spectrum access method based on attention mechanism and echo state network

Through the Att-DESN framework with enhanced attention, the problems of scarcity and fragmentation of spectrum resources in 6G networks are solved, efficient and intelligent spectrum management is achieved, user collision rate is reduced, and QoS scores are improved, which is suitable for complex network environments.

CN120263321AActive Publication Date: 2025-07-04HANGZHOU DIANZI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510736531.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

The existing technology has problems of scarcity and fragmentation of spectrum resources in 6G networks. The traditional DSA method has problems such as low spectrum allocation efficiency, poor adaptability, insufficient robustness and limited network compatibility in the future, especially in high dynamic environments, Q value overestimation and decision-making accuracy are limited.

Method used

The dual echo state network (Att-DESN) framework based on attention enhancement is adopted, combined with additive attention mechanism and ESN, a mixed reward function and QoS scoring model are designed, and spectrum characteristics are processed through dual deep Q network (DDQN) to achieve efficient and intelligent distribution of spectrum management.

Benefits of technology

Significantly reduce the collision rate between secondary users and primary users, improve service quality (QoS) scores, adapt to heterogeneous network environments, improve system throughput and convergence speed, and is suitable for terrestrial communications, satellite communications and Internet of Things scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120263321A_ABST
    Figure CN120263321A_ABST
Patent Text Reader

Abstract

The invention discloses a dynamic spectrum access method based on an attention mechanism and an echo state network, and the method comprises the steps: firstly constructing a 6G space-ground integrated network system model and a three-dimensional network topology model, and defining a state space and an action space; secondly, a Bahdanau attention mechanism is introduced into a DDQN framework, and a mixed framework with time sequence modeling capacity is constructed in combination with the short-term memory characteristic of the ESN; and then designing a reward function considering spectrum efficiency and PU protection, and evaluating network QoS through a weighted multi-index model. And finally, on the basis of a Markov state space and an action space, dynamic spectrum features are processed through a hybrid architecture, a reward function and a QoS evaluation mechanism are combined, and an offline-online two-stage optimization strategy is adopted, so that autonomous dynamic spectrum access of the space-ground integrated network is completed. According to the invention, heterogeneous requirements of a ground network, a satellite network and the Internet of Things are compatible, and remarkable performance improvement is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - field of wireless communication technology and artificial intelligence technology, and specifically relates to a dynamic spectrum access (DSA) method based on attention - enhanced deep reinforcement learning, especially a dynamic spectrum allocation method based on an attention - enhanced deep Q - network (DQN). Background Art

[0002] The next - generation wireless network faces the challenges of scarce and fragmented spectrum resources. With the rapid growth of IoT devices and the research and deployment of 6G networks, the wireless communication network is facing an increasingly tight spectrum resource problem. It is predicted that by 2030, the number of globally connected devices will exceed 200 billion. The existing static spectrum allocation method can no longer meet the explosive growth of spectrum demand. As a key solution to solve the spectrum scarcity problem, DSA technology allows SUs to intelligently share licensed spectrum resources without interfering with PUs.

[0003] Traditional DSA methods are mainly divided into three categories: one is the distributed coordination method based on game theory, which has low decision - making efficiency and is difficult to converge; the second is the perception - based heuristic algorithm, which has poor adaptability to environmental dynamic changes; the third is the early reinforcement learning method, such as Q - learning, which has certain learning ability but faces the problem of state - space explosion. In recent years, deep reinforcement learning (DRL) has been introduced into the DSA field, such as DQN and its variants. Although it partially solves the problem of high - dimensional state processing, there are still the following technical bottlenecks: the over - estimation problem of Q - value leads to policy deviation; the decision - making reliability is insufficient in a partially observable environment; the adaptability to heterogeneous user scenarios is poor; the convergence speed is slow and the training stability is insufficient.

[0004] In addition, after the introduction of millimeter - wave and terahertz bands in 6G networks, signal attenuation intensifies and environmental interference is significant, and a more intelligent spectrum allocation strategy is urgently needed. For example, when ultra - reliable low - latency communication coexists with enhanced mobile broadband, DSA needs to improve the system throughput while ensuring high reliability. The performance of existing methods in a partially observable environment still needs to be improved, especially in high - dynamic scenarios where satellite communication coexists with ground networks. Although the existing patent CN117896027A uses a DQN framework to achieve spectrum allocation, there are still two key defects in practical applications: firstly, this scheme uses a traditional DQN structure and fails to effectively solve the over - estimation problem of Q - value in a high - dynamic environment; secondly, its state representation completely depends on the original observation data and lacks a selective attention mechanism for spectrum features, resulting in limited decision - making accuracy in a partially observable scenario. Therefore, an innovative solution that can solve the above problems simultaneously is urgently needed. Summary of the Invention

[0005] The objective of the present invention is to provide a DSA method and system based on an attention-enhanced dual echo state network (Att-DESN) to solve the problems of low spectrum allocation efficiency, poor adaptability, insufficient robustness, and limited future network compatibility in the prior art. The present invention combines an attention mechanism with a double deep Q-network (DDQN) framework of an echo state network (ESN) to solve the problems of scarce and fragmented spectrum resources in 6G and Internet of Things (IoT) networks, and achieve efficient and intelligent spectrum management.

[0006] The present invention is applicable to future heterogeneous network environments, including scenarios where terrestrial communication, satellite communication, and a large number of IoT devices coexist, and can significantly reduce the collision rate between secondary users and primary users (SU-PU), improve the quality of service (QoS) score, and provide key technical support for dynamic spectrum sharing in 6G networks.

[0007] The present invention has achieved a technological breakthrough through the following core innovations:

[0008] (1) Innovatively combines an additive attention mechanism with an ESN, dynamically allocates the importance of spectrum features through attention weights, and simultaneously utilizes the time-series processing ability of the ESN to capture the dynamic characteristics of the spectrum;

[0009] (2) Proposes a hybrid reward function and a QoS scoring model to optimize the reliable and high-QoS communication objectives on the premise of ensuring the hard constraint of the collision rate of PUs;

[0010] (3) Aiming at the characteristics of 6G space-ground integrated networks, expands the path loss model and introduces a target network soft synchronization mechanism, significantly improving the policy convergence and robustness under high-attenuation channels.

[0011] The technical solution adopted by the present invention is: a DSA method based on an attention-enhanced dual deep echo state network, including the following steps:

[0012] Step 1, construct a 6G space-ground integrated network system model: establish a distributed DSA scenario including N primary users (PUs) and M secondary users (SUs), where each PU exclusively occupies a wireless channel, and the SUs, as autonomous agents, dynamically adjust their access strategies by sensing the channel state; construct a three-dimensional network topology model, including a three-layer heterogeneous architecture of low-earth orbit satellite nodes, terrestrial base stations, and mobile terminals.

[0013] Step 2, design the state space and action space: define the observation state space of the SUs as a Markov process of the PU activity state, and the action space as a channel selection or waiting decision;

[0014] Step 3, Integrating the Attention Mechanism and ESN: Introduce the Bahdanau attention mechanism into the DDQN framework to dynamically focus on the key features of the input data, enhancing the model's ability to handle time-varying and noisy spectral environments; at the same time, combine the short-term memory characteristics of ESN to construct a hybrid architecture with temporal modeling capabilities;

[0015] Step 4, Optimizing the Reward Function and QoS Scoring: Design a reward function that takes into account both spectral efficiency and PU protection, and evaluate the network QoS through a weighted multi-index model;

[0016] Step 5, Training and Deployment: Based on the Markov state space and discrete action space constructed in Step 2, process the dynamic spectral features through the attention-enhanced dual ESN architecture in Step 3, combine the multi-objective reward function and QoS evaluation mechanism designed in Step 4, and adopt an offline-online two-stage optimization strategy: in the offline stage, use experience replay and dual network synchronous training to achieve policy convergence, and in the online stage, deploy a lightweight model to achieve real-time spectrum decision-making, and finally complete the autonomous dynamic spectrum access in the space-ground integrated network environment.

[0017] Preferably, Step 3 uses an attention-enhanced dual echo state network to achieve feature extraction and decision optimization of the spectrum state. The Bahdanau attention mechanism processes the input observations through a three-layer structure: First, the spectral features are linearly transformed and activated by the tanh function to obtain the hidden representation , where the learnable parameter is initialized using a normal distribution, where ensures the sparsity of feature selection; subsequently, the feature importance is calculated through the scoring function and the attention weights are generated using the softmax function; the enhanced state retains the original information through a residual connection, where represents element-wise multiplication. This enhanced state is input into the dual echo state network, and its dynamic equation is . The input weight matrix projects the attention-enhanced state to the reservoir layer, and the reservoir weight matrix is , and its spectral radius is set to 0.6 to ensure the echo state property. The output weight matrix is a trainable output layer used to generate Q-value estimates. Under the structure of DDQN, action selection and evaluation are implemented through separate and networks, and the target value , where, represents the immediate reward, is the discount factor, is the target network, is the evaluation network, is the next state, is the next action, are the parameters of the evaluation network, are the parameters of the target network. Finally, the output of the double echo state network with attention enhancement is , where is the current state-action pair.

[0018] Preferably, in step 4, the reward function is designed in four cases according to the SU behavior: the reward when successfully accessing is proportional to the logarithm of the channel capacity: ; a penalty related to the interference level is imposed when there is a collision between SUs , and the function uses the X logarithm form to map the interference value to the interval to ensure that the penalty intensity is proportional to the interference degree. The fixed penalty when there is a collision between the PU and the SU is ; the reward is 0 when no channel is selected. Among them, represents the secondary user at time selects the channel the signal-to-interference-plus-noise ratio (SINR), which is determined by the channel gain , the primary user interference power , the interference power between secondary users and the thermal noise power jointly, and the specific calculation formula is . Among them is the transmission power of the primary user, represents the path loss between the primary user and the secondary user, represents the channel state (1 for occupied, 0 for idle). is the channel bandwidth, is the power spectral density of the thermal noise.

[0019] Preferably, in step 4, the QoS score is achieved by weighted combination of multiple key performance indicators, and its expression is , where each component is normalized. represents the normalized average access success rate of the SU, and the calculation method is , where is the total number of access attempts, represents the th attempt is successful. and respectively represent the average collision rates after normalization with respect to PU and SU, and the calculation methods are similar. The calculation formula for the PU collision rate is represents the collision event with PU. represents the average throughput after normalization, and its original value is determined by the channel bandwidth , signal-to-interference-plus-noise ratio , noise loss factor and the training batch duration jointly, and the calculation formula is , where is the SINR of SU on the channel . represents the normalized value of the long-term reward, and its original value is obtained by discounting and accumulating the immediate rewards, and the calculation formula is , where is the discount factor, is the immediate reward of SU at time , is the time range. The weight coefficient satisfies , and can be adjusted according to the specific application scenario. For example, in the ultra-reliable low-latency communication scenario, (the weight of the SU-PU collision rate) can be increased to preferentially guarantee the quality of service of PU.

[0020] Preferably, in step 5, the training process adopts the optimization strategy of deep reinforcement learning, and the specific implementation includes three key components: the experience replay buffer used to store the transition samples , and the sample correlation is broken and the training efficiency is improved by randomly sampling a small batch of data; the exploration rate adopts a linear decay strategy, starting from the initial value and gradually decaying to . This design enables the agent to fully explore the environment in the initial stage of training and gradually tend to utilize the learned strategies in the later stage; the learning rate is set to during the training process, and the discount factor . These hyperparameters have been verified by a large number of experiments to obtain the best performance. The capacity of the experience replay buffer is set to , and each time training, samples are uniformly sampled from for network parameter update, and this mechanism effectively alleviates the training instability problem caused by the data sequence correlation.

[0021] Preferably, in step 5, the online deployment system implements the following functions: The real-time spectrum sensing module collects channel state observations at a fixed period (typical value is 10 ms). These observations include the occupancy status of each channel (0 indicates idle, 1 indicates occupied by the primary user), SINR estimation values and other key information; After receiving the observations, the attention mechanism module first calculates the hidden representation , where is a learnable parameter matrix, is the hidden layer dimension, is the -th feature component of the observation; Subsequently, through attention weight calculation and normalization to obtain the enhanced state ; The decision engine inputs the enhanced state into the trained Att-DESN network, evaluates the value of each action through the Q function , and selects the optimal channel access action , where represents the network parameters; The system maintenance mechanism includes real-time monitoring of channel characteristic change indicators . When exceeds the preset threshold (typical value is 0.3), the model fine-tuning process is automatically triggered. This threshold is determined based on historical data statistics to ensure that the system always adapts to the current wireless environment.

[0022] Preferably, in step 5, the training environment is configured as: Based on the python3.12 and Anaconda environment, using libraries such as tensorflow, numpy, and os for configuration.

[0023] The present invention has the following characteristics and beneficial effects:

[0024] High spectrum utilization rate and low collision rate: Through the attention-enhanced DDESN framework, the present invention achieves significant performance improvement in complex scenarios such as 6G networks and the Internet of Things.

[0025] Cross-network architecture compatibility: The unique channel modeling method is compatible with the heterogeneous requirements of terrestrial networks, satellite networks, and the Internet of Things. By integrating the WINNER II path loss model and the unique atmospheric attenuation term of satellite communication, the system can simultaneously handle spectrum management in millimeter wave (30 - 300 GHz), terahertz (0.1 - 10 THz), and traditional UHF frequency bands.

[0026] Fast convergence and computational efficiency: Combining the advantages of ESN and DDQN, the training convergence speed is increased by 57.1% compared with traditional Q-learning. The fixed random reservoir design of ESN reduces the computational complexity to , and the sparse processing of the attention mechanism reduces the communication overhead by more than 30%, making it suitable for deployment on IoT devices with limited computing resources.

[0027] Scalable system architecture: The modularly designed reward function and QoS scoring system support flexible parameter adjustment by modifying the weight coefficient It can be adapted to different application scenarios such as industrial Internet of Things (low latency priority) and smart city (high reliability priority). System throughput model Supports bandwidth configurations from 36MHz to 1GHz, meeting diverse needs from narrowband IoT to enhanced mobile broadband. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] The technical solution of the present invention is described in detail below in conjunction with the accompanying drawings. It should be noted that the accompanying drawings are only used to exemplarily describe the embodiments of the present invention so that those skilled in the art can better understand the technical solution of the present invention. Without departing from the core idea of ​​the present invention, those skilled in the art can obtain other implementation methods based on the accompanying drawings.

[0029] Figure 1 A DSA system training flow chart of an embodiment of the present invention;

[0030] Figure 2 A schematic diagram of a ground-sky integrated system model according to an embodiment of the present invention;

[0031] Figure 3 Schematic diagram of the three-dimensional spatial distribution of PU and SU in an embodiment of the present invention.

[0032] The specific implementation details of the above figures can be further understood in conjunction with the embodiments of the present invention. Those skilled in the art can adjust the system configurations and parameters shown in the figures according to actual application scenarios, but these adjustments are within the protection scope of the present invention. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0034] like Figure 1As shown in the figure, the training process of the Att-DESN model of the present invention includes steps such as environmental perception, attention enhancement, network training, and model update: First, the agent weights the historical observations through the attention mechanism to generate enhanced observations; then the observations are input into the echo state network for state expansion and stored in the experience replay pool together with actions and rewards; subsequently, data is sampled from the experience replay pool to update the DDQN-ESN network parameters; finally, the agent interacts with entities such as low-orbit satellites, PU / SU terminals, etc. in the Markov spectrum environment to obtain new observations and rewards, realizing the iterative optimization of the strategy.

[0035] In the description of the present invention, it should be understood that the orientation or positional relationship indicated by terms such as "center", "longitudinal", "lateral", "upper", "lower", "front", "rear", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and thus cannot be understood as a limitation of the present invention. In addition, terms such as "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise specified, the meaning of "a plurality" is two or more.

[0036] In the description of the present invention, the term "attention mechanism" refers to the Bahdanau attention mechanism, which is used to dynamically weight input features; the term "spectrum sensing" refers to obtaining the channel state information of primary users and secondary users through wireless signal detection technology; the term "echo state network" refers to a recurrent neural network structure with fixed random weights, which is used to process time-series data.

[0037] Embodiment 1

[0038] This embodiment provides a DSA method for future 6G networks, and its system architecture is as Figure 2 shown. Figure 2 This is a schematic diagram of the system model of the present invention, in which the low-orbit satellite senses the spectrum environment, the PU transmitter sends an ideal signal, the SU transmitter generates a potential interference signal, the PU terminal and the SU receiver receive and feedback the channel state, and the agent outputs an action to access the available channel according to the sensing result and the attention-enhanced representation, and obtains corresponding rewards according to the PU / SU conflict situation.

[0039] This method combines the Bahdanau attention mechanism with DDQN and ESN to achieve efficient spectrum resource allocation. The specific implementation steps are as follows:

[0040] Step 1: System initialization phase.

[0041] In the system initialization phase, network parameters are first set, including the reference distance , path loss exponent , channel gain coefficient and . Meanwhile, communication parameters are configured, the PU transmission power dBm, the SU transmission power dBm, the channel bandwidth is MHz, the attention mechanism parameters are initialized to follow the normal distribution, and a replay buffer D with a capacity of 300 pieces of empirical data is established.

[0042] Step 2: Spectrum sensing and state representation.

[0043] During the spectrum sensing and state representation process, the SU monitors the occupancy status of channels in real time through a wireless receiver. A system state space is constructed, where represents the occupancy status of the th channel. Considering a 20% sensing error probability , the environment is modeled as a partially observable Markov decision process.

[0044] Step 3: Attention enhancement processing.

[0045] In the attention enhancement processing stage, the Bahdanau attention mechanism is applied to the observed state . First, the hidden state is calculated, then the attention score is calculated, and the attention weight is obtained through the softmax function. Finally, the enhanced state representation is output.

[0046] Step 4: Decision-making and action execution.

[0047] When making a decision and executing an action, is input into the DDQN-ESN network, where the ESN part uses for state update. An action is selected through the -greedy strategy, and the optimal action is selected with probability , and an available action is randomly selected with probability . The channel access action is executed, where 0 means staying idle, Indicates which channel is selected.

[0048] Step 5: Reward calculation and learning stage.

[0049] In the reward calculation and learning stage, the immediate reward is calculated according to four situations : When successfully accessing, the reward is proportional to the logarithm of the channel capacity; when there is a conflict between SUs, a penalty is imposed according to the interference level; when there is a conflict between PU and SU, a fixed penalty -C3 is imposed; when no channel is selected, the reward is 0. Store the experience into the replay buffer D, and sample batch data from D to update the network parameters.

[0050] Step 6: Performance evaluation stage.

[0051] When evaluating the performance, calculate the QoS score, comprehensively considering the access success rate , PU conflict rate , SU conflict rate , normalized throughput and long-term cumulative reward and other indicators, and use the weight coefficient to perform weighted summation to obtain the final QoS score.

[0052] Figure 3 is a schematic diagram of the spatial distribution of the PU transmitter, PU receiver, SU transmitter, and SU receiver in the X-Y-Z plane in the simulation scenario. This figure shows the random distribution of the transmitters (Transmitter) and receivers (Receiver) of PUs and SUs within a space of 50km × 50km × 500km. A total of 8 PUs and 4 SUs are arranged in the test scenario. Circles and triangles in the figure represent transmitters and receivers respectively, and the coordinates are in kilometers, which are used to construct the path loss model and evaluate the spectrum allocation performance at different positions.

[0053] Specifically, the technical effects of this embodiment are verified through simulation experiments: in the test scenario containing 6 PUs and 2 SUs, compared with the current advanced DEQN method, the present invention reduces the SU-PU conflict rate by 70% and improves the QoS score by 7.2%. In the more complex 8PU-4SU scenario, it still maintains the advantage of reducing the conflict rate by 33%, and at the same time, the convergence speed is improved by 57.1% compared with the Q-learning method. These improvements are mainly due to the fact that the DDQN structure effectively alleviates the problem of overestimation of Q values, the attention mechanism enhances the ability to extract key features, and the ESN architecture improves the efficiency of processing time-series data.

[0054] Table 1

[0055] The proposed Att-DESN framework of the present invention can be further extended and applied to other wireless resource management scenarios, including but not limited to fields such as power control, user access selection, and network slice resource allocation. The Att-DESN training process is shown in Table 1. By adjusting the feature extraction method of the attention mechanism and the design of the reward function, this framework can adapt to the dynamic decision-making requirements in different communication scenarios, demonstrating good generality and scalability. In particular, in intelligent communication systems that need to handle partially observable environments, the method proposed in the present invention has significant technical advantages and application potential.

[0056] It should be understood that the above embodiments are only the preferred embodiments of the present invention. For those skilled in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.

[0057] Embodiment 2

[0058] This embodiment provides a DSA system based on Att-DESN, including one or more processors and a storage module.

[0059] The storage module is used to store program codes and model parameters. When the program is executed by the processor, the processor can call the pre-trained Att-DESN model to make DSA decisions based on the real-time simulated wireless environment observation data, and update the model parameters to achieve adaptive optimization.

[0060] This embodiment also provides a computer-readable storage medium, which stores a computer program that can implement the above DSA method when executed by the processor. The storage medium includes media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.

[0061] The embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings, but the present invention is not limited to the described embodiments. For those skilled in the art, without departing from the principle and spirit of the present invention, various changes, modifications, substitutions, and variations to these embodiments including components still fall within the protection scope of the present invention.

Claims

1. A dynamic spectrum access method based on an attention mechanism and an echo state network, characterized in that It includes the following steps: Step 1: Construct a 6G space-ground integrated network system model; Step 2: Define the observation state space of the secondary user SU in the 6G space-ground integrated network system model as a Markov process of the primary user PU's activity state, and the action space as channel selection or waiting decision; Step 3: Introduce the Bahdanau attention mechanism in the DDQN framework, dynamically focus on the key features of the input data, and at the same time combine the short-term memory characteristics of the ESN to construct a hybrid architecture with time series modeling ability; Step 4: Design a reward function that takes into account both spectrum efficiency and PU protection, and evaluate the network QoS through a weighted multi-index model; Step 5: Based on the Markov state space and action space, process the dynamic spectrum features through the hybrid architecture, combine the reward function and the QoS evaluation mechanism, and adopt an offline-online two-stage optimization strategy to complete the autonomous dynamic spectrum access in the space-ground integrated network environment.

2. The dynamic spectrum access method based on the attention mechanism and the echo state network according to claim 1, wherein The 6G space-ground integrated network system model includes a distributed DSA scenario with N primary users PU and M secondary users SU, where each PU exclusively occupies a wireless channel, and the SU, as an autonomous intelligent agent, dynamically adjusts the access strategy by sensing the channel state; The three-dimensional network topology model of the 6G space-ground integrated network system includes a three-layer heterogeneous architecture of low-earth orbit satellite nodes, ground base stations, and mobile terminals.

3. The dynamic spectrum access method based on the attention mechanism and the echo state network according to claim 2, characterized in that The specific implementation process of Step 3 is as follows: The Bahdanau attention mechanism processes the input observations through a three-layer structure: first, the spectral features are represented as a hidden representation obtained through linear transformation and activation by the tanh function ; subsequently, the feature importance is calculated through a scoring function and the attention weights are generated using the softmax function ; the enhanced state is the product of the spectral features and the corresponding attention weights, and the original information is retained through a residual connection; Enhance the state input double echo state network to obtain a hidden representation , which serves as the input data under the DDQN framework.

4. The dynamic spectrum access method based on the attention mechanism and the echo state network according to claim 3, wherein The reward function described in step 4 is designed in four cases according to the SU behavior: the reward at successful access is proportional to the logarithm of the channel capacity: , where represents the signal-to-interference-plus-noise ratio (SINR) of the secondary user k when selecting channel c at time t; a penalty related to the interference level is imposed during collisions between SUs , and the function adopts the X logarithm form to map the interference value to the interval ; the fixed penalty during collisions between PUs and SUs is ; the reward is 0 when no channel is selected.

5. The dynamic spectrum access method based on the attention mechanism and the echo state network according to claim 4, wherein In step 4, the QoS is achieved by weighted combination of multiple performance metrics, and the expression is , where each component has been normalized; represents the average access success rate of the SU after normalization; and respectively represent the average collision rates of the PU and the SU after normalization; represents the average throughput after normalization; represents the normalized value of the reward, and the weight coefficient satisfies .

6. The dynamic spectrum access method based on the attention mechanism and the echo state network according to claim 5, characterized in that, The specific implementation process of Step 5 is as follows: The training process adopts an optimization strategy of deep reinforcement learning, and an experience replay buffer is used to store transition samples , where represents the state, represents the action, represents the immediate reward, is the next state; the exploration rate adopts a linear decay strategy; the value of each action is evaluated through the Q function to select the optimal channel access action , where represents the network parameters.

Citation Information

Patent Citations

  • Distributed dynamic spectrum allocation method and device based on deep reinforcement learning

    CN117896027A

  • Dynamic spectrum and power control method based on cyclic Flash reinforcement learning

    CN118870549A

  • Echo cancellation algorithm

    CN119007737A