A Gibbs Sampling-Inspired Dynamic Spectrum Access Method for Covert Communication

By employing a rejection mechanism inspired by Gibbs sampling, the channel selection strategy is optimized, solving the problem of insufficient covert communication capabilities in existing technologies and achieving more efficient spectrum utilization and covert communication performance.

CN120568336BActive Publication Date: 2026-04-03TONGJI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-23
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing reinforcement learning-based dynamic spectrum access methods lack a proper balance between adaptability, exploration efficiency, and communication overhead, and fail to effectively improve covert communication capabilities. In particular, users are easily detected and tracked when facing adaptive attacks.

Method used

We employ a Gibbs sampling-inspired rejection mechanism to probabilistically deny access to overused channels, constructing a probabilistic rejection-based DSA strategy. Combined with the Q-Learning framework, we optimize channel selection to reduce the probability of detection by attackers and improve covert communication performance.

Benefits of technology

It effectively reduces the probability of users continuously accessing the same channel, improves the balance and concealment of spectrum utilization, enhances users' covert communication capabilities, and reduces the possibility of detection and tracking by attackers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120568336B_ABST
    Figure CN120568336B_ABST
Patent Text Reader

Abstract

This invention proposes a Gibbs sampling-inspired Dynamic Spectrum Access (DSA) method for covert communication. First, the covert communication metric is defined as the KL divergence between the user channel access distribution and a uniform distribution. The lower the KL divergence, the more difficult it is for an attacker to detect or attack the transmission channel. Second, inspired by Gibbs sampling, a rejection mechanism is proposed to deny access to overused channels. This encourages users to dynamically access multiple channels, reducing the probability of users continuously accessing the same channel. Finally, the rejection mechanism is combined with a Q-Learning reinforcement learning algorithm to form a novel aperiodic frequency-hopping DSA framework for covert communication. By continuously adjusting channel selection, the performance of covert communication is improved. Simulation results show that, compared with the standard Q-Learning algorithm, under different user densities, the proposed frequency-hopping DSA method significantly reduces the KL divergence while maintaining good channel capacity, thus enhancing the performance of covert communication.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of wireless communication and networks, and specifically relates to cognitive radio technology based on reinforcement learning. Background Technology

[0002] The continued rapid growth in global wireless technology usage has brought new challenges to ensuring secure and private data transmission. Due to the open and broadcast nature of wireless channels, communication systems are inherently vulnerable to various forms of interference and malicious attacks. In particular, attackers can eavesdrop on ongoing transmissions or interfere by sending high-powered signals into the communication channel. Such attacks severely reduce the signal-to-interference-plus-noise ratio (SINR) of legitimate signals in normal communication, leading to data transmission failures or data that cannot be understood by the receiving end. These risks highlight the growing need for covert communication technologies to reduce the likelihood of detection or attack during spectrum access. This problem is particularly acute in areas with high communication security requirements, such as military operations, financial systems, and government communications, where ensuring the integrity, confidentiality, and availability of data is paramount. Therefore, protecting wireless transmissions from attackers—including those aiming to eavesdrop, inject jamming, or tamper with signals—remains a fundamental challenge for wireless security. With the evolution of increasingly complex radio services, the need for adaptive intelligent covert communication technologies is more urgent than ever before.

[0003] Frequency hopping (FH) technology has been widely adopted to reduce the risk of attack. A FH system first communicates with the receiver to predefine a channel switching pattern. Users transmit information on rapidly switching channels, thus reducing the possibility of attackers continuously interrupting transmissions on a specific frequency band. FH systems have proven effective against non-adaptive attackers who detect and attack transmission channels based on static or pre-defined patterns without real-time feedback from the radio environment. However, with the development of spectrum sensing and adaptive signal processing technologies, attackers have developed adaptive attack strategies capable of detecting, tracking, and adapting to the target user's channel selection strategy. Due to the periodic nature of FH system channel selection, it typically lacks the sensitivity required to respond to rapidly changing radio environments.

[0004] Meanwhile, the complexity and variability of the radio environment have also spurred the development of Dynamic Spectrum Access (DSA) technology. DSA enables users to observe and respond to changes in the radio environment in real time, promoting efficient and flexible spectrum utilization. Although DSA was initially designed to alleviate spectrum shortages, its core advantages—spectrum awareness and adaptive decision-making—make it particularly suitable for covert communications. Among various DSA methods, reinforcement learning (RL) has become one of the most effective and widely adopted solutions. RL-based DSA users can autonomously track optimal channel selection strategies by continuously interacting with the radio environment and learning from feedback. In recent years, the integration of covert communications and intelligent learning has become a hot topic. Many studies are exploring how to utilize RL to enhance the performance of covert communications in dynamic radio environments.

[0005] However, existing methods often involve trade-offs between adaptability, exploration efficiency, and communication overhead. Furthermore, many of these methods primarily focus on throughput or interference avoidance, neglecting the importance of enhancing the user's own covert communication capabilities—that is, reducing the likelihood of detection or tracking by attackers. These limitations necessitate a more balanced and effective anti-attack strategy that promotes diverse channel access without relying on frequent inter-user communication or user movement. Summary of the Invention

[0006] To address the inherent limitations of existing RL-based DSA methods, this invention proposes a Gibbs sampling-inspired dynamic spectrum access method for covert communication, which fundamentally enhances the flexibility and covertness of spectrum access strategies.

[0007] Specifically, this invention designs a rejection mechanism inspired by Gibbs sampling, which probabilistically rejects access to overused channels. This method evaluates the channel access probability of each channel and selectively rejects overused channels, guiding users to more diverse and unpredictable channel selection outcomes.

[0008] This method not only reduces the risk of users being easily detected by attackers due to prolonged access to the same channel, but also promotes a more balanced use of available spectrum. Therefore, users can achieve stronger covert communication performance without frequent communication with the receiver or the need for user relocation to avoid attacker interference.

[0009] Technical solution

[0010] A Gibbs sampling-inspired method for covert communication dynamic spectrum access includes the following steps:

[0011] Step 1: Each SU creates its own Q-Table. In each time slot, when the user senses the channel state, it looks up the Q-value vector corresponding to that state in the Q-Table, which is the Q-value corresponding to all optional actions.

[0012] Step 2: SU statistically analyzes the action selection results over a period of time prior to the current time slot, calculates the historical channel access frequency, and designs a probability-based rejection mechanism based on the Gibbs sampling concept.

[0013] Step 3: Integrate the rejection mechanism into the Q-Learning reinforcement learning framework to construct a probabilistic rejection-based DSA policy and make the final channel selection decision.

[0014] Through the above steps and multiple access iterations, the optimal spectrum access strategy for cognitive users is obtained, at which point the covert communication capability reaches its optimal level.

[0015] Beneficial effects

[0016] Inspired by Gibbs sampling, this invention designs a rejection mechanism for channel access decisions in dynamic spectrum access. Unlike the standard Q-Learning algorithm's strategy of always selecting the optimal channel, this rejection mechanism probabilistically rejects access to the optimal channel, reducing the probability of continuous access to the same channel. This prevents users from continuously accessing the same channel, effectively reducing the probability of being detected and attacked by attackers, and improving the user's ability to conduct covert communication. Attached Figure Description

[0017] Figure 1 This is a schematic diagram of the processing flow of the method of the present invention;

[0018] Figure 2 This is a schematic diagram of a dynamic spectrum access scenario according to an embodiment of the present invention;

[0019] Figure 3 This is a diagram of a covert communication DSA framework based on a rejection mechanism and Q-Learning, according to an embodiment of the present invention.

[0020] Figure 4 This is an embodiment of the present invention. =4 and Performance comparison chart of the algorithm when =8 ((a)) =6(b) =8(c) =12). Detailed Implementation

[0021] The present invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0022] This invention relates to a Gibbs sampling-inspired method for covert communication dynamic spectrum access. The complete process for users to perform dynamic spectrum access is as follows: Figure 1 As shown.

[0023] In dynamic spectrum access scenarios, there are Primary Users (PUs) Secondary Users (SUs) and Each independent channel, such as Figure 2 As shown.

[0024] First, the user creates their own Q-Table (Q-value state action table, where the state corresponds to different channel state vectors and the action represents selecting different channels). In each time slot, when the user perceives the channel state, they look up the Q-value vector corresponding to that state in the Q-Table, which is the Q-value corresponding to all available actions.

[0025] At the same time, users will statistically analyze the action selection results in the previous period of the current time slot, calculate the historical channel access frequency to approximate the current channel access probability, and design a probability-based rejection mechanism based on the idea of ​​Gibbs sampling.

[0026] Gibbs sampling is a Markov chain Monte Carlo method that extracts samples from a high-dimensional joint distribution. Its core idea is to avoid the curse of dimensionality caused by directly sampling the joint distribution by iteratively conditionally sampling each variable. Inspired by the Gibbs sampling mechanism, this invention introduces a rejection mechanism into the action selection strategy. Specifically: (1) Drawing on the idea of ​​Gibbs sampling “sampling each variable condition in turn”, a condition selection process is designed that proceeds sequentially, starting from the action with the highest priority and gradually trying whether to adopt the action; (2) Similar to constructing a condition distribution based on other variables in Gibbs sampling, in each step of selection, the probability of the action being rejected is calculated based on the historical access pattern of the channel. This probabilistic rejection mechanism is analogous to the condition sampling of a variable under the joint distribution in Gibbs sampling; (3) When selecting an action, if the current high-priority channel is rejected due to overuse, the system tries the next alternative action. This method retains the Markov chain property of Gibbs sampling “updating according to variable conditions”, and introduces perturbation while controlling the search range of space, thereby avoiding getting trapped in a suboptimal local solution.

[0027] Then, the user will jointly decide on the final channel selection based on the rejection mechanism and the Q-value ranking in the Q-value vector.

[0028] Finally, the access point receives the user's message and, based on the reward function, feeds back the actual channel capacity obtained by the user. The user then updates the Q-Table based on the feedback result. The next time slot is then sensed and accessed, thus completing the user's dynamic spectrum access.

[0029] A Gibbs sampling-inspired method for covert communication dynamic spectrum access, the specific steps of which are as follows:

[0030] Step 1: Each SU creates its own Q-Table. In each time slot, when the user senses the channel state, it looks up the Q-value vector corresponding to that state in the Q-Table, which is the Q-value corresponding to all optional actions.

[0031] Specifically, the steps are as follows: Step (11): Each SU constructs its own Q-Table (Q-value state action table), where the state of the Q-Table corresponds to different channel state vectors, and the action represents the selection of different channels.

[0032] The Q-Table is used to calculate the action selection for channel access and to update the Q value after the reward is obtained.

[0033] Examples of Q-Tables include... Figure 3 The state is represented as Action is represented as .

[0034] Step (12): In the time slot Obtain channel state awareness results Then, each SU looks up the corresponding Q-value vector in its own Q-Table.

[0035]

[0036] in The last element represents the total number of accessible channels. This indicates the Q value corresponding to the action "not connected". This represents the Q-value of the m-th action. The corresponding Q value for selecting the m-th channel.

[0037] Then sort the Q values ​​in the Q-value vector:

[0038]

[0039] in This indicates the action to be performed after sorting by Q value. This represents the action corresponding to the maximum Q value. This represents the action corresponding to the minimum Q value.

[0040] Step 2: SU statistically analyzes the action selection results over a period of time prior to the current time slot, calculates the historical channel access frequency, and designs a probability-based rejection mechanism based on the Gibbs sampling concept.

[0041] Specifically, the steps are as follows: Step (21): Each SU maintains a channel access probability vector. For the channel The channel access probability is defined as the probability of channel access occurring within a certain period of time prior to the current time slot. Number of accesses and within that time period The ratio of the total number of accesses to each channel.

[0042] The ideal channel access probability follows a uniform distribution, which is expressed as:

[0043]

[0044] A sliding window is introduced to calculate the channel access probability, with the sliding window size for each SU being... ,in Sliding window size and channel access probability update period The proportionality constant is 5 in this example. Channel access probability per Each time slot is periodically updated. SU for channel The channel access probability is defined as

[0045]

[0046] in Indicates SU in time slot Access Channel Specifically, if SU is in the time slot Select access channel ,So ;otherwise, .

[0047] Each SU independently updates its own channel access probability vector, which serves as the basis for subsequent channel selection decisions.

[0048] Step (22): The difference in channel access probability between two consecutive candidate actions is defined as...

[0049]

[0050] in Indicates action The probability of choosing, Indicates action The probability of selection. At the current moment. The calculation refers to the aforementioned formula ( ).parameter The larger the value, the more likely the SU will choose to act. The probability is higher than the choice of action The probability is higher. SU is more likely to fall into the trap of accessing only a single channel, making it easier for attackers to detect and attack.

[0051] Define the rejection probability as

[0052] ,

[0053] Where parameters Defined as the level of covert communication performance, this parameter is set to 3 in this example. A smaller value means a higher probability of channel access and a greater probability of the action being rejected (corresponding to a higher probability of action rejection for a higher Q value). This enhances SU's covert communication capabilities and avoids excessive access to the same channel.

[0054] if For the last candidate action Its rejection probability is

[0055]

[0056] Here, the rejection probability of the last candidate action ensures that the action is only likely to be rejected when its channel access probability is greater than the uniform distribution threshold.

[0057] Introduce a uniformly distributed random variable ,when If the current action is selected, then the algorithm moves to the next candidate action and re-evaluates whether to reject it. If all actions are rejected, SU will select the action "Do Not Access".

[0058] Step 3: Integrate the rejection mechanism into the Q-Learning reinforcement learning framework to construct a probabilistic rejection-based DSA policy to determine the final channel selection. For example... Figure 3 As shown. Includes the following steps:

[0059] Step (31): In the time slot The actual channel state is ,in Indicates the first The actual state of each channel. Indicates channel Not occupied Indicates channel It is already in use. First, SU performs channel sensing, and the sensing result is... ,in Indicates SU to the first The perception results for each channel. In real-world scenarios, perception is often imperfect and prone to errors. Assume the probability of a perception error is... The probability that SU correctly perceives the channel state is

[0060]

[0061] Perception Results As input to the Q-Learning algorithm framework.

[0062] Step (32): SU based on the perception results Select the corresponding Q-value vector, and then select the action according to the Q-value sorting and probability rejection mechanism described in steps 1 and 2.

[0063] Step (33): After the SU accesses the channel according to the selected action, the environment provides a corresponding reward based on the reward function. The reward function is expressed as follows:

[0064]

[0065] in This represents the penalty value when SU ​​collides with PU. This refers to the channel capacity when the SU is accessing the network.

[0066] Step (34): The Q-value matrix is ​​determined according to the reward function. Update the Q value:

[0067]

[0068] in, For learning rate, . The discount factor represents the degree to which future rewards influence current decisions.

[0069] Step (35): After multiple iterations, the KL divergence converges. The KL divergence is defined as the channel access probability distribution. With uniform distribution The degree of proximity. It is expressed as...

[0070]

[0071] The main contributions of the above technical solution are as follows: 1) KL divergence is defined as a measure of concealment: This invention introduces the KL divergence between the user channel access distribution and the uniform distribution as a measure to quantify concealment performance. The lower the KL divergence, the more balanced and unpredictable the channel access behavior, making it more difficult for attackers to detect or attack. 2) Gibbs sampling-inspired rejection mechanism: This mechanism probabilistically rejects access to overused channels. This mechanism encourages users to dynamically explore multiple channels, reducing long dwell times on any single channel. 3) Frequency-hopping DSA framework for covert communication: This invention integrates the rejection mechanism into the Q-Learning reinforcement learning framework, forming a novel DSA strategy based on non-periodic frequency hopping. This strategy improves covert communication performance by reducing KL divergence.

[0072] like Figure 4 for =4 and When the number of users is 8, the algorithm of this invention and the standard Q-Learning algorithm are compared at different numbers of users. A comparison of covert communication performance (Random mode with completely random access), among which... Figure 4 (a) =6; Figure 4 (b) =8; Figure 4 (c) =12. Simulation results show that, compared with the standard Q-Learning algorithm, the algorithm of this invention can reduce the KL divergence by 17.8%, 18.9% and 15.4% respectively, that is, it has better covert communication performance.

[0073] The above description is merely a description of preferred embodiments of this application and is not intended to limit the scope of this application in any way. Any changes or modifications made by those skilled in the art based on the above-disclosed technical content should be considered as equivalent and valid embodiments and fall within the scope of protection of the technical solution of this application.

Claims

1. A Gibbs sampling-inspired method for covert communication dynamic spectrum access, characterized in that, Includes the following steps: Step 1: Each SU creates its own Q-Table. In each time slot, when the user senses the channel state, it looks up the Q-value vector corresponding to that state in the Q-Table, which is the Q-value corresponding to all possible actions. Step 2: SU statistically analyzes the action selection results over a period of time prior to the current time slot, calculates the historical channel access frequency, and designs a probability-based rejection mechanism based on the Gibbs sampling concept. Step 3: Integrate the rejection mechanism into the Q-Learning reinforcement learning framework to construct a probabilistic rejection-based DSA strategy to determine the final channel selection; Through the above steps, after multiple access iterations, the optimal spectrum access strategy for cognitive users is obtained, at which point the covert communication capability reaches its optimal level. Step 2 is as follows: Step (21): Each SU maintains a channel access probability vector For the channel The channel access probability is defined as the probability of channel access occurring within a certain period of time prior to the current time slot. Number of accesses and within that time period The ratio of the total number of access attempts to each channel; A sliding window is introduced to calculate the channel access probability, with the sliding window size for each SU being... ,in Sliding window size and channel access probability update period Proportional constant; Channel access probability per Each time slot is periodically updated; Time slot SU for channel The channel access probability is defined as in Indicates SU in time slot Access Channel Specifically, if SU is in the time slot Select access channel ,So ; otherwise, ; Each SU independently updates its own channel access probability vector, which serves as the basis for subsequent channel selection decisions. Step (22): The difference in channel access probability between two consecutive candidate actions is defined as... in Indicates action The probability of choosing, Indicates action The probability of choosing; Define the rejection probability as , Where parameters Defined as a level of covert communication performance; if For the last candidate action Its rejection probability is Introduce a uniformly distributed random variable ,when If the current action is selected, the algorithm moves to the next candidate action and decides whether to reject it again. If all actions are rejected, the algorithm will select the action "Do not access". Step 3 specifically involves: Step (31): In the time slot The actual channel state is ,in Indicates the first The actual state of each channel Indicates channel Not occupied Indicates channel It has been occupied; First, SU performs channel sensing, and the sensing result is: ,in Indicates SU to the first Perception results of each channel; perception results As input to the Q-Learning algorithm framework; Step (32): SU based on the perception results Select the corresponding Q-value vector, and then select the action based on the Q-value sorting and probability rejection mechanism; Step (33): After the SU accesses the channel according to the selected action, the environment feeds back the corresponding reward according to the reward function; the reward function is expressed as follows: in This represents the penalty value when the SU collides with the PU; The channel capacity when the SU is accessed; Step (34): The Q-value matrix is ​​determined according to the reward function. Update the Q value: in, For learning rate, ; The discount factor represents the degree to which future rewards influence current decisions; Step (35): After multiple iterations, the KL divergence converges; the KL divergence is defined as the channel access probability distribution. With uniform distribution The degree of closeness is expressed as 。 2. The Gibbs sampling-inspired covert communication dynamic spectrum access method according to claim 1, characterized in that, Step 1 is as follows: Step (11): Each SU constructs its own Q-value state action table Q-Table, where the state of the Q-Table corresponds to a different channel state vector, and the action represents the selection of a different channel; Step (12): In the time slot Obtain channel state awareness results Then, each SU looks up the corresponding Q-value vector in its own Q-Table. in The last element represents the total number of accessible channels. This indicates the Q value corresponding to the action "do not connect"; This represents the Q-value of the m-th action. The corresponding Q value for selecting the m-th channel; Then sort the Q values ​​in the Q-value vector: in This indicates the instructions for the corresponding actions after sorting by Q value; This represents the action corresponding to the maximum Q value. This represents the action corresponding to the minimum Q value.

3. The Gibbs sampling-inspired covert communication dynamic spectrum access method according to claim 1, characterized in that, The value is 5.

4. The Gibbs sampling-inspired covert communication dynamic spectrum access method according to claim 1, characterized in that, The value is 3.

Citation Information

Patent Citations

  • Sparse channel estimation method based on approximate sampling and underwater acoustic communication system

    CN118041727A

  • Fast convergence dynamic induction spectrum access method based on LSTM and Q-Learning fusion

    CN118764114A