A knowledge-driven anti-interference spectrum access method and system
By constructing a spectrum access model under complex interference environments and combining it with a Markov decision process based on prior knowledge, channel selection and transmission duration are optimized. This solves the problem of slow convergence speed in traditional reinforcement learning algorithms, achieves fast and effective spectrum access, and improves communication performance and spectrum utilization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NAT UNIV OF DEFENSE TECH
- Filing Date
- 2022-12-16
- Publication Date
- 2026-07-31
AI Technical Summary
In complex interference environments, traditional reinforcement learning algorithms converge slowly, making it difficult to help user equipment quickly obtain effective spectrum access strategies, resulting in degraded wireless communication performance and low spectrum utilization.
A spectrum access model under interference environment is constructed, and the channel selection and transmission duration decision problem is modeled as a Markov decision process. An anti-interference spectrum access algorithm based on prior knowledge is used to guide the exploration process of reinforcement learning by optimizing channel selection and transmission duration, initializing the Q-table using prior state, and designing a composite reward function.
It improved the algorithm's convergence speed, increased communication throughput, enhanced spectrum utilization, reduced false alarm and missed detection probabilities, and optimized the transmission strategy of user equipment.
Smart Images

Figure CN115988658B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of spectrum access technology, and in particular to a knowledge-driven anti-interference spectrum access method and system. Background Technology
[0002] With the advent of the 5G mobile communication era, a new round of technological revolution and industrial transformation is developing in depth, bringing great convenience to life. However, emerging industries, such as cloud computing, VR / AR, and autonomous driving, are making the problems of spectrum resource shortage and low spectrum utilization increasingly prominent. Although high-frequency bands, such as millimeter wave bands and terahertz bands, are being developed and applied to 5G and 6G, traditional static spectrum allocation methods still cannot meet the rapidly growing communication demands.
[0003] To address the aforementioned issues, an artificial intelligence-based approach has been proposed to develop an efficient spectrum sharing scheme. However, this measure results in the openness of the spectrum in wireless communication networks, making user equipment signals highly susceptible to malicious interference, thus reducing the effectiveness and reliability of wireless communication. Therefore, ensuring user equipment signal transmission performance and improving spectrum utilization in complex interference environments has significant theoretical and practical value.
[0004] Dynamic spectrum access technology is a key technology for achieving efficient utilization of spectrum resources in wireless communication systems. It possesses the ability to detect spectrum holes and select idle spectrum for transmission, making it an efficient spectrum sharing scheme. The complex and variable electromagnetic spectrum under interference environments poses challenges to dynamic spectrum access for user equipment. Reinforcement learning, which continuously learns from its environment, can solve the real-time spectrum decision-making problem in dynamic interference environments. However, in complex interference electromagnetic environments, traditional reinforcement learning suffers from slow convergence speed and long convergence periods, making it difficult to help user equipment quickly obtain an effective access strategy in the initial stage. Summary of the Invention
[0005] The purpose of this invention is to provide a knowledge-driven anti-interference spectrum access method and system that can ensure the signal transmission performance of user equipment and improve spectrum utilization in complex interference environments.
[0006] To achieve the above objectives, the present invention provides the following solution:
[0007] In a first aspect, the present invention provides a knowledge-driven anti-interference spectrum access method, comprising:
[0008] A spectrum access model under interference conditions is constructed, and the channel selection and transmission duration decision problems of user equipment are determined based on the spectrum access model under interference conditions. The spectrum access model under interference conditions is constructed based on a wireless communication system. The wireless communication system includes a user equipment and an jammer. The user equipment includes at least a transmitter and a receiver.
[0009] The channel selection and transmission duration decision problem is modeled as a Markov decision process.
[0010] Determine the prior state, and determine the anti-interference spectrum access algorithm that combines prior knowledge based on the prior state;
[0011] The Markov decision process is solved using the anti-interference spectrum access algorithm that incorporates prior knowledge to obtain the transmission channel and transmission duration for the user equipment's dynamic access spectrum.
[0012] The transmission channel and transmission duration of the spectrum specifically include:
[0013] Secondly, the present invention provides a knowledge-driven anti-interference spectrum access system, comprising:
[0014] The optimization problem determination module is used to construct a spectrum access model under interference conditions and determine the channel selection and transmission duration decision problem of user equipment based on the spectrum access model under interference conditions. The spectrum access model under interference conditions is constructed based on a wireless communication system. The wireless communication system includes a user equipment and an jammer. The user equipment includes at least a transmitter and a receiver.
[0015] A Markov decision process construction module is used to model the channel selection and transmission duration decision problem as a Markov decision process.
[0016] The algorithm determination module is used to determine the prior state and determine the anti-interference spectrum access algorithm based on the prior state and prior knowledge.
[0017] The transmission channel and transmission duration determination module is used to solve the Markov decision process using the anti-interference spectrum access algorithm that combines prior knowledge, so as to obtain the transmission channel and transmission duration of the user equipment's dynamic access spectrum.
[0018] According to specific embodiments provided by the present invention, the present invention discloses the following technical effects:
[0019] To address the dynamic spectrum access decision-making problem under interference environments, this invention leverages the advantage of prior knowledge in improving the convergence speed of reinforcement learning algorithms, providing a knowledge-driven anti-interference spectrum access method and system. Simulation results show that this invention not only improves algorithm convergence speed and shortens learning time, but also increases communication throughput and improves spectrum utilization. The main work includes:
[0020] First, considering the impact of frequent channel switching on communication throughput, this invention simultaneously optimizes channel selection and transmission duration, modeling the joint decision-making problem of channel selection and transmission duration as a Markov decision process.
[0021] Second, for the dynamic spectrum access decision-making problem under interference environments, a dynamic spectrum access anti-interference algorithm combining prior knowledge is proposed. The changing trend of interference is used as prior knowledge, four prior states are defined and their priorities are determined, and the Q-table is initialized. Simultaneously, a composite reward function is designed based on the prior states to guide the reinforcement learning exploration process.
[0022] In the simulation section, the algorithm proposed in this invention is compared with traditional reinforcement learning without prior knowledge and conventional prior knowledge methods in terms of performance metrics such as average throughput, false alarm probability, and false miss probability. Furthermore, this invention also analyzes the impact of different numbers of prior states on performance metrics. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0024] Figure 1 This is a flowchart illustrating the anti-interference spectrum access method provided in an embodiment of the present invention.
[0025] Figure 2 This is a structural diagram of a spectrum access model under interference conditions provided in an embodiment of the present invention;
[0026] Figure 3 A time-slot structure diagram provided for an embodiment of the present invention;
[0027] Figure 4 The jamming effect diagram of the jammer on 5 different channels provided in the embodiment of the present invention;
[0028] Figure 5 A block diagram illustrating the principle of reinforcement learning provided in an embodiment of the present invention;
[0029] Figure 6A block diagram illustrating the principle of reinforcement learning based on prior knowledge, provided for embodiments of the present invention;
[0030] Figure 7 A comparison chart of performance metrics for four different spectrum access algorithms provided in embodiments of the present invention; Figure 7 (a) shows the average throughput curves for different dynamic spectrum access algorithms; Figure 7 (b) False alarm probability curves for different dynamic spectrum access algorithms; Figure 7 (c) Missed detection probability curves for different dynamic spectrum access algorithms;
[0031] Figure 8 A comparison chart showing the impact of different numbers of prior states on performance indicators provided in embodiments of the present invention; Figure 8 (a) Comparison of the impact of different numbers of prior states on average throughput; Figure 8 (b) Comparison of the impact of different numbers of prior states on the false alarm probability; Figure 8 (c) Comparison of the impact of different numbers of prior states on the probability of missed detection. Detailed Implementation
[0032] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0033] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0034] Example 1
[0035] like Figure 1 As shown in the figure, this embodiment of the invention provides a knowledge-driven anti-interference spectrum access method, which includes the following steps.
[0036] Step 100: Construct a spectrum access model under interference conditions, and determine the channel selection and transmission duration decision problem of user equipment based on the spectrum access model under interference conditions; the spectrum access model under interference conditions is constructed based on the wireless communication system.
[0037] This invention addresses the dynamic spectrum access problem for user equipment (UEs) under interference environments. During dynamic spectrum access, UEs are susceptible to malicious interference from the external environment. This problem is mitigated by learning historical information and real-time status of the complex electromagnetic environment.
[0038] The spectrum access model under interference conditions described in this embodiment of the invention is as follows: Figure 2 As shown, the constructed wireless communication system has one user equipment (containing a transmitter and a receiver) and one jammer. The wireless communication system has N available non-overlapping channels with a bandwidth of BHz. The transmitter's transmit power is P, and the jammer's interference power is J, which varies with time.
[0039] When the user equipment selects any one channel for data transmission, the transmitter transmits data to the receiver through the communication link. The receiver determines the control link feedback decision information based on the received data and the execution of the intelligent decision algorithm, and transmits it to the transmitter. The jammer is used to simultaneously interfere with N channels in the wireless communication system and to block the normal communication of the user equipment by releasing interference signals.
[0040] Both the communication link and the interference link utilize Rayleigh fading channels. The channel gain is related to the transmission distance, and their relationship can be expressed as:
[0041] φ=d -δ (1);
[0042] Where d represents the transmission distance, δ represents the path fading factor, and φ represents the channel gain μ or interference gain β of the user equipment. This embodiment of the invention assumes that the transmit power P of the user equipment is variable during communication, and that the user equipment is subject to interference from the external environment, leading to data loss and channel fading. Therefore, the signal-to-noise ratio (SINR) of the user equipment on channel i can be expressed as:
[0043]
[0044] Where i represents the channel (i∈N), μ represents the channel gain from transmitter to receiver, P represents the transmitter's transmit power, N0 represents the channel background noise power, β represents the interference gain from jammer to receiver, and J i J represents the interference power of the channel. th This represents the interference threshold, where τ is an indicator function representing the level of interference experienced by the user equipment. i >J th When the signal is 0, it indicates that the user equipment is being interfered with; otherwise, the user equipment is not being interfered with.
[0045] The specific definitions are as follows:
[0046]
[0047] Meanwhile, to facilitate computation and intelligent decision-making, this embodiment of the invention divides the transmission duration discretely into several time slot units T of equal length. fThe transmission time slot set is defined as t = {1, 2, ..., T}. The time slot structure is as follows: Figure 3 As shown, each time slot unit T f It consists of four parts, T s Part of this involves the channel awareness and selection phase, where user equipment senses the interference status of the channel and selects a channel with better communication quality. d Part of this is the data transmission phase, where the user equipment communicates data based on the sensing results from the previous phase. ACK Part of the process is the data transmission confirmation phase, T l Part of this is the learning phase, where the communication throughput and reward value for this time slot are calculated, and decision information is updated. Therefore, the communication throughput obtained by the user equipment in each time slot unit is expressed as:
[0048]
[0049] The dynamic spectrum access process designed in this embodiment of the invention includes several time slots t (t∈T), and the communication throughput obtained by the user equipment in each dynamic spectrum access is expressed as:
[0050]
[0051] Figure 4 The image shows the jamming effect of the jammer on five channels. The heatmap changes from dark to light colors, indicating that the jamming effect on user equipment is more pronounced. In actual communication, user equipment needs to find channel interference holes to transmit data for an appropriate duration. However, the transmission duration cannot be too long. When the transmission duration is too long, it may collide with interference, leading to data loss and reducing the communication quality of the user equipment. Conversely, if the transmission duration is too short, the user equipment will frequently switch channels, resulting in huge energy consumption.
[0052] Therefore, the goal of the spectrum access model constructed in this embodiment of the invention is to quickly find the optimal transmission strategy in an unknown interference environment, that is, to select the optimal transmission channel and transmission duration, reduce the probability of user equipment being interfered with, and maximize communication throughput.
[0053] Step 200: Model the channel selection and transmission duration decision problem as a Markov Decision Process (MDP).
[0054] Q-learning is a value-based reinforcement learning algorithm. Its advantage lies in combining temporal difference algorithms with optimization control theory, and performing offline learning through temporal difference methods. The reinforcement learning framework contains three typical elements: (1) the agent's state S. n (2) The agent's action an (3) The reward value r obtained n The Q-table constructed using reinforcement learning can reflect the quality of decisions, and the principle is as follows: Figure 5 As shown, in the discrete time series 0, 1, ..., n, the agent is initially in state S. n Action a is obtained through intelligent decision-making. n The AI will receive a reward value r when it performs this action. n Then, the agent updates the Q-table based on the reward value of the current state. The agent repeats this process in a loop, constantly providing feedback to adjust the strategy, and finally obtains the optimal decision.
[0055] During reinforcement learning, the agent selects an action based on the current Q-value at each step. If the user device always chooses the action corresponding to the maximum Q-value, the algorithm will get stuck in a local optimum. To avoid this, this embodiment of the invention employs a greedy strategy algorithm for updating. This algorithm balances exploration and exploitation, allowing the agent to learn a globally optimal strategy through continuous learning. The agent selects the action 'a' with the maximum Q-value with a probability of 1-ε, and randomly selects action A with a probability of ε. Specifically:
[0056]
[0057] When describing an MDP, three elements are used: state S, action A, and reward R. Their specific meanings in this paper are as follows:
[0058] (1) State space S
[0059] In this embodiment of the invention, the state of the user equipment in time slot k is defined as... This represents the interference power experienced by the user equipment in time slots k-2 and k-1, β represents the interference gain from the jammer to the receiver, k∈K, and K represents the total simulation time for the user's dynamic access settings. Therefore, in the case of n channels, the state space is as follows:
[0060]
[0061] (2) Action Space A
[0062] In time slot k, the user equipment will be in S k The action selected in the state is defined as:
[0063] A k =[n k ,t k (8);
[0064] Where, n k Indicates that the user equipment is in state S kThe selected transmission channel, t k Indicates that the user equipment is in state S k Select the transmission duration below.
[0065] (3) Reward function R
[0066] User equipment in S k Execute action A in the state k You will receive a corresponding reward value R. k Because user equipment in S k Execute action A in the state k Will choose t k The data transmission duration, where t k (t∈T) contains several time slots, and the interference experienced by the user equipment in each time slot is different. Therefore, the reward received by the user equipment in each time slot is different, as shown below:
[0067]
[0068] The formula introduces the information transmission rate C = Blog2(1 + SINR) as the evaluation of reward, a1 and a2 represent two dimensions of behavior, and P t Let represent the transmit power of the user equipment in any time slot during a single channel access process. Therefore, the immediate return of the user equipment in the dynamic spectrum access process in time slot k can be expressed as:
[0069]
[0070] In this process, a single access procedure comprises several time slots t (t∈T), therefore, the formula... This represents the benefit of a user equipment accessing a channel once; on the other hand, ηmJ th J represents the cost of interference to user equipment, η represents the cost factor, and J represents the cost of interference to user equipment. th Let represent the interference threshold, and m represent the number of times a user equipment (UE) experiences interference during a single channel access process. Therefore, the reward value R in the entire formula comprises two parts: firstly, it penalizes the UE for experiencing interference during channel access, and secondly, it encourages the UE to use longer time slots during a single channel access process to reduce the number of handovers.
[0071] Step 300: Determine the prior state and determine the anti-interference spectrum access algorithm based on the prior state and prior knowledge.
[0072] The state space S defined in this embodiment of the invention reflects the multi-channel interference effect of two time slots. Based on the characteristics of the interference time series, the changing trend of the interference time series can be found from the state space. Combined with the interference threshold, four channel access scenarios can be identified. These four access scenarios are defined as the prior state S. *Prior state S * The channel quality of the access channel selected by the user equipment at a future time is predicted based on the prior state S. * For Q-table initialization, a channel selection and transmission duration optimization algorithm based on reinforcement learning is proposed, and it is combined with prior state S. * The algorithm employs a "decision-feedback-adjustment" approach, which involves continuously adjusting and improving its transmission strategy by evaluating the quality of actions in each state. The principle block diagram is shown in Figure 6.
[0073] Prior state S * The specific definitions are as follows:
[0074]
[0075]
[0076] Among them, the prior state In This indicates that the interference time series on channel n shows a decreasing trend. This indicates that the user equipment was not disturbed at time k-1. Therefore, combining the two, we can conclude that the user equipment will not be disturbed for a long time in the future, and the user equipment is in a priori state. This is the most suitable location for dynamic access.
[0077] Prior state In This indicates that the interference time series on channel n shows an upward trend. This indicates that the user equipment was not disturbed at time k-1. Therefore, combining the two relationships, we can conclude that the user equipment will not be disturbed in the near future, and the user equipment is in a priori state. This is more suitable for dynamic access.
[0078] Prior state In This indicates that the interference time series on channel n shows a decreasing trend. This indicates that the user equipment was interfered with at time k-1. Therefore, combining the two, we can conclude that the user equipment will be interfered with in the near future, and the user equipment is in a priori state. Dynamic access is not suitable for this application.
[0079] Prior state In This indicates that the interference time series on channel n shows an upward trend. This indicates that the user equipment was interfered with at time k-1. Therefore, combining the two, we can conclude that the user equipment will be interfered with for a long time to come, and the user equipment is in a priori state. Dynamic access is not suitable for this application.
[0080] Based on the above analysis, the prior state This illustrates four scenarios in the dynamic spectrum access process for user equipment, including the prior state. It is the optimal access state of the user equipment, the prior state. This is the worst-case access state for the user equipment; therefore, four prior states with varying priorities are defined. The Q-table of the user equipment is initialized according to the priority of the prior states.
[0081] Step 400: Solve the Markov decision process using the anti-interference spectrum access algorithm that combines prior knowledge to obtain the transmission channel and transmission duration of the user equipment's dynamic access spectrum.
[0082] When a user equipment completes a dynamic spectrum access process, it will check which prior state S this access process conforms to. * Provide additional rewards R to user devices:
[0083]
[0084] Where ω is the benefit enhancement coefficient. This is the discount factor.
[0085] Therefore, when a user equipment completes a dynamic spectrum access process, combined with the prior state S * Reshape the payoff function R:
[0086]
[0087] During the learning process, the user equipment continuously interacts with the environment to explore the changing patterns of interference, thereby obtaining the optimal transmission strategy. The optimization objective of this embodiment is to optimize the user equipment's transmission strategy π, such that Q under the current strategy... π To maximize the long-term cumulative utility of the system, the Q-function is iteratively updated as follows:
[0088]
[0089] Where α represents the learning rate and γ represents the decay factor of future reward values.
[0090] Step 400 is operated as follows:
[0091] Obtain the initial channel state s through spectrum sensing 0 The user equipment randomly selects a transmission channel and transmission duration for data communication.
[0092] Set the initial time slot k = 0; Loop: for k = 0 to k max .
[0093] User equipment in s k Select transmission channel n under state k , with t k The duration of data transmission. Determine the user equipment's duration in seconds. k What type of prior state S does the current state belong to? * The compound reward R obtained from this action is calculated according to formula (13). T .
[0094] User equipment in s k Under the given state, the transmission channel n for the next time step is obtained using an ε-greedy strategy according to formula (6). k+1 and transmission duration t k+1 .
[0095] Update the Q value according to formula (14).
[0096] Update status s k+1 →s k .
[0097] When k = k + 1, the loop ends.
[0098] The dynamic spectrum access algorithm incorporating prior knowledge can be summarized as follows: the system defines four prior states based on the state space and interference threshold, and initializes the Q-table of the user equipment; it iterates cyclically according to the principle of maximizing reward value through reinforcement learning to optimize channel selection and transmission duration. The core steps of the algorithm are shown in Table 1.
[0099] Table 1. Dynamic Spectrum Access Anti-interference Algorithm Based on Prior Knowledge
[0100]
[0101]
[0102] This invention simulates and tests a dynamic spectrum access algorithm under interference conditions in the MATLAB platform environment, comparing four different spectrum access schemes based on evaluation metrics (false alarm probability, missed detection probability, and average throughput). Scheme 1 is a traditional single-action reinforcement learning algorithm, using a greedy algorithm to decide the channel selection action and a random action for transmission duration selection; Scheme 2 is a traditional dual-action reinforcement learning algorithm, using a greedy algorithm to decide the channel selection and transmission duration action; Scheme 3 is a reinforcement learning algorithm based on conventional prior knowledge, using prior knowledge to initialize the Q-table and accelerate the algorithm's convergence speed; Scheme 4 is the reinforcement learning algorithm proposed in this paper that combines prior knowledge, using prior knowledge to define four prior states, initializing the Q-table, and designing a composite reward function using the prior states to accelerate the algorithm's convergence speed; it is assumed that the interference power J of the jammer in the communication system ranges from 2.5W to 10W, and the interference period T = 40 time slots. Considering the power waste problem, this paper assumes that when J < J th At that time, the transmission power P = J, if J > J th Transmit power P = J th To avoid the algorithm getting trapped in local optima, ε is set to a value that decreases uniformly with the simulation time slots. The simulation time slots are set to 40000. The specific communication system simulation parameters and reinforcement learning parameters are shown in Tables 2 and 3.
[0103] Table 2 Simulation parameters and reinforcement learning parameters for the communication system
[0104]
[0105]
[0106] Table 3 Reward Parameters for Different Prior States
[0107]
[0108] Figure 7 The figure shows a comparison of the performance metrics of four different spectrum access algorithms. Figure 7 Figure (a) shows a comparison of the average throughput of four different spectrum access schemes. As can be seen from the figure, the traditional reinforcement learning algorithm without prior knowledge, the conventional prior knowledge algorithm, and the proposed reinforcement learning algorithm incorporating prior knowledge begin to converge around the 35,000th, 25,000th, and 20,000th time slots, respectively. Compared with the comparative algorithms, the proposed algorithm can improve the convergence speed by approximately 42% and 16%, respectively, demonstrating a significant advantage. Furthermore, the algorithm proposed in this embodiment achieves better average throughput in the initial stage of the access process and, during the convergence stage, significantly outperforms the traditional reinforcement learning algorithm without prior knowledge and slightly outperforms the conventional prior knowledge algorithm.
[0109] Figure 7(b) shows a comparison of false alarm probabilities for four different spectrum access schemes. As can be seen from the figure, the comparative algorithm begins to converge around the 35,000th time slot, while the proposed algorithm begins to converge around the 15,000th time slot, representing a 57% improvement in convergence speed. After convergence, the false alarm probabilities of the comparative algorithm stabilize at 0.2, 0.1, and 0.045, respectively, while the proposed algorithm stabilizes at around 0.025, representing reductions of 17.5%, 7.5%, and 2%, respectively, effectively improving spectrum utilization. Figure 7 (c) shows a comparison of the missed detection probabilities for four different spectrum access schemes. As can be seen from the figure, the algorithm proposed in this embodiment of the invention has a lower missed detection probability in the initial stage of the access process and a significant advantage in convergence speed. During the convergence stage, the missed detection probability obtained by the algorithm proposed in this embodiment of the invention is superior to the comparative algorithms, reducing the probability of interference to user equipment and improving average throughput.
[0110] Secondly, this invention also investigates the impact of the number of prior states on performance indicators. While keeping the dynamic spectrum access model and algorithm parameters unchanged under interference conditions, Figure 8 (a) The impact of different numbers of prior states on average throughput was studied. As shown in the figure, with the increase in the number of prior states possessed by the user equipment, the average throughput gradually increases and the convergence speed significantly accelerates. This indicates that a larger number of prior states is more beneficial to improving the algorithm's convergence performance. On the other hand, the convergence phase in the figure shows that prior states... The average throughput under these conditions is significantly lower than that with prior states. The results of the other three cases indicate that the prior state This is the optimal access situation for user equipment. When user equipment has prior state s1 initialization, it can not only help user equipment converge quickly, but also improve the average throughput. Conversely, when user equipment does not have prior state s1 initialization, user equipment is easily interfered with when accessing the channel, the convergence speed is slow, and the communication throughput is reduced.
[0111] Figure 8 (b) and (c) investigated the impact of different numbers of prior states on the false alarm probability and the false alarm probability. As can be seen from the figure, when the user equipment has prior states... In the case where the user equipment has more prior states, the convergence speed of the false alarm probability and the missed detection probability obtained during the access process is faster and gradually decreases. Meanwhile, when the user equipment has no prior states... In the case of s1, the probability of false alarms and missed detections obtained by the user equipment is significantly higher than in the other three cases, and it takes more time to reach convergence. This indicates that the prior state s1 can help the user equipment select a suitable channel and transmission time slot, reducing the probability of false alarms and missed detections during the channel access process.
[0112] Example 2
[0113] To implement the method corresponding to Embodiment 1 above and achieve the corresponding functions and technical effects, a knowledge-driven anti-interference spectrum access system is provided below. This system includes:
[0114] The optimization problem determination module is used to construct a spectrum access model under interference conditions and determine the channel selection and transmission duration decision problem of user equipment based on the spectrum access model under interference conditions. The spectrum access model under interference conditions is constructed based on a wireless communication system. The wireless communication system includes a user equipment and an jammer. The user equipment includes at least a transmitter and a receiver.
[0115] The Markov decision process construction module is used to model the channel selection and transmission duration decision problem as a Markov decision process.
[0116] The algorithm determination module is used to determine the prior state and determine the anti-interference spectrum access algorithm based on the prior state and prior knowledge.
[0117] The transmission channel and transmission duration determination module is used to solve the Markov decision process using the anti-interference spectrum access algorithm that combines prior knowledge, so as to obtain the transmission channel and transmission duration of the user equipment's dynamic access spectrum.
[0118] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple; relevant parts can be referred to the method section.
[0119] This document uses specific examples to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. Furthermore, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A knowledge-driven anti-jamming spectrum access method, characterized in that, include: Construct a spectrum access model under interference conditions, and determine the channel selection and transmission duration decision problem of user equipment based on the spectrum access model under interference conditions; The spectrum access model under interference conditions is constructed based on a wireless communication system; the wireless communication system includes a user equipment and an jammer; the user equipment includes at least a transmitter and a receiver; The channel selection and transmission duration decision problem is modeled as a Markov decision process. Determine the prior state, and determine the anti-interference spectrum access algorithm that combines prior knowledge based on the prior state; The expression for the prior state is: ; ; Among them, the prior state In This indicates that the interference time series on channel n shows a decreasing trend. Indicates that the user equipment is Since the device has not been disturbed at any time, combining the two factors, we can conclude that the user equipment will not be disturbed for a long period of time in the future, and the user equipment is in a priori state. This is the most suitable environment for dynamic access; Prior state In This indicates that the interference time series on channel n shows an upward trend. Indicates that the user equipment is Since the device has not been disturbed at any time, combining the two factors, we can conclude that the user equipment will not be disturbed in the near future, and the user equipment is in a priori state. The following is more suitable for dynamic access; Prior state In This indicates that the interference time series on channel n shows a decreasing trend. Indicates that the user equipment is Since the device is constantly subject to interference, combining the two factors, we can conclude that the user equipment will be subject to interference in the near future, and the user equipment is in a priori state. Dynamic access is not suitable in the following situations; Prior state In This indicates that the interference time series on channel n shows an upward trend. Indicates that the user equipment is Since the user equipment is constantly subject to interference, combining the two factors, we can conclude that the user equipment will be subject to interference for a long period of time in the future, and the user equipment is in a priori state. Dynamic access is not suitable in the following situations; The Markov decision process is solved using the anti-interference spectrum access algorithm incorporating prior knowledge to obtain the transmission channel and transmission duration for dynamic access spectrum of user equipment, specifically including: Obtaining the initial channel state through spectrum sensing The user equipment randomly selects a transmission channel and transmission duration for data communication; Set the initial time slot Loop: for to ; User equipment is in status Select transmission channel ,by The transmission duration of the transmitted data determines the status of the user equipment. Subordinate prior states ; According to the formula Calculate the compound reward obtained from this action. ;in, For the benefit enhancement coefficient, This is the discount factor; This represents the cost of interference with user equipment. Represents the cost factor. Indicates the interference threshold. This indicates the number of times a user equipment is interfered with during a single access channel session; Indicates channel gain; User equipment is in status Below, according to the formula by Greedy strategy to obtain the transmission channel at the next time step and transmission duration ; According to the formula Update the Q value; where, Indicates the learning rate, A decay factor representing the future reward value; Update status ; The loop ends.
2. The knowledge-driven anti-interference spectrum access method according to claim 1, characterized in that, The wireless communication system has N available non-overlapping channels with a bandwidth of BHz. The transmitter's transmission power is P, and the jammer's interference power is J, which varies with the time period. When the user equipment selects any one channel for data transmission, the transmitter transmits data to the receiver through the communication link. The receiver determines the control link feedback decision information based on the received data and the execution of the intelligent decision algorithm, and transmits it to the transmitter. The jammer is used to simultaneously interfere with N channels in the wireless communication system and to block the normal communication of the user equipment by releasing interference signals.
3. The knowledge-driven anti-interference spectrum access method according to claim 2, characterized in that, In the wireless communication system, both the communication link and the interference link use Rayleigh fading channels.
4. The knowledge-driven anti-interference spectrum access method according to claim 2, characterized in that, Before performing the step of modeling the channel selection and transmission duration decision problem as a Markov decision process, the method further includes: dividing the transmission duration into several time slot units of equal length, and determining the communication throughput obtained by the user equipment each time it accesses the spectrum based on the time slot units. The formula for calculating the communication throughput is: ; in, Indicates that the user equipment is in any time slot The communication throughput obtained within, , Indicates channel bandwidth. Indicates that the user equipment is in the channel Signal-to-noise ratio; This refers to the channel sensing and channel selection phase. Indicates the data transmission stage; This indicates the data transmission confirmation phase; This indicates the learning stage.
5. The knowledge-driven anti-interference spectrum access method according to claim 4, characterized in that, The Markov decision process includes: state space The state space S represents the user equipment in... The power of interference received within the time slot; Action space A; in Time slots will keep user equipment in a state The action to be selected is defined as: ; in, Indicates the user equipment is in a certain state. Select the transmission channel below. Indicates the user equipment is in a certain state. Select the transmission duration below; Indicates in Status of time-slot user equipment; reward function The reward function includes penalties for user equipment accessing the channel being interfered with, and rewards for user equipment to use longer time slots during a single channel access process to reduce the number of handovers.
6. The knowledge-driven anti-interference spectrum access method according to claim 4, characterized in that, The prior state is used to predict the channel quality of the access channel selected by the user equipment at future times; the prior state determines the anti-interference spectrum access algorithm that combines prior knowledge, specifically including: Determine the priority of the prior states; The Q-table is initialized using the prior states with priority, resulting in a channel selection and transmission duration optimization algorithm based on reinforcement learning; the anti-interference spectrum access algorithm is a channel selection and transmission duration optimization algorithm based on reinforcement learning.
7. A knowledge-driven anti-interference spectrum access system, characterized in that, include: The optimization problem determination module is used to construct a spectrum access model under interference conditions and determine the channel selection and transmission duration decision problem of user equipment based on the spectrum access model under interference conditions. The spectrum access model under interference conditions is constructed based on a wireless communication system. The wireless communication system includes a user equipment and an jammer. The user equipment includes at least a transmitter and a receiver. A Markov decision process construction module is used to model the channel selection and transmission duration decision problem as a Markov decision process. An algorithm determination module is used to determine a priori states and, based on the priori states, determine an anti-interference spectrum access algorithm incorporating prior knowledge; the expression for the priori states is: ; ; Among them, the prior state In This indicates that the interference time series on channel n shows a decreasing trend. Indicates that the user equipment is Since the device has not been disturbed at any time, combining the two factors, we can conclude that the user equipment will not be disturbed for a long period of time in the future, and the user equipment is in a priori state. This is the most suitable environment for dynamic access; Prior state In This indicates that the interference time series on channel n shows an upward trend. Indicates that the user equipment is Since the device has not been disturbed at any time, combining the two factors, we can conclude that the user equipment will not be disturbed in the near future, and the user equipment is in a priori state. The following is more suitable for dynamic access; Prior state In This indicates that the interference time series on channel n shows a decreasing trend. Indicates that the user equipment is Since the device is constantly subject to interference, combining the two factors, we can conclude that the user equipment will be subject to interference in the near future, and the user equipment is in a priori state. Dynamic access is not suitable in the following situations; Prior state In This indicates that the interference time series on channel n shows an upward trend. Indicates that the user equipment is Since the user equipment is constantly subject to interference, combining the two factors, we can conclude that the user equipment will be subject to interference for a long period of time in the future, and the user equipment is in a priori state. Dynamic access is not suitable in the following situations; The transmission channel and transmission duration determination module is used to solve the Markov decision process using the anti-interference spectrum access algorithm that incorporates prior knowledge, to obtain the transmission channel and transmission duration of the user equipment's dynamic access spectrum. Specifically, it includes: obtaining the initial state of the channel through spectrum sensing. The user equipment randomly selects a transmission channel and transmission duration for data communication; Set the initial time slot Loop: for to ; User equipment is in status Select transmission channel ,by The transmission duration of the transmitted data determines the status of the user equipment. Subordinate prior states ; According to the formula Calculate the compound reward obtained from this action. ;in, For the benefit enhancement coefficient, This is the discount factor; This represents the cost of interference with user equipment. Represents the cost factor. Indicates the interference threshold. This indicates the number of times a user equipment is interfered with during a single access channel session; Indicates channel gain; User equipment is in status Below, according to the formula by Greedy strategy to obtain the transmission channel at the next time step and transmission duration ; According to the formula Update the Q value; where, Indicates the learning rate, A decay factor representing the future reward value; Update status ; The loop ends.