A dynamic spectrum access method, apparatus, device and medium
By constructing a dynamic spectrum access time slot structure in the Internet of Vehicles and using a continuous long short-term memory dual-duel deep Q network to optimize the spectrum access strategy, the problems of high energy consumption and access failure under the traditional static spectrum allocation method are solved, and efficient and low-energy spectrum access and communication are achieved.
Patent Information
- Application Number
- CN202510308379.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2026-01-23
- Estimated Expiration
- 2045-03-17
AI Technical Summary
In the Internet of Vehicles (IoV), traditional static spectrum allocation methods cannot adapt to rapidly changing communication needs, resulting in low spectrum resource utilization, increased energy consumption, frequent access failures, and low channel utilization, which affect communication quality and security.
A dynamic spectrum access method is adopted. By determining the dynamic spectrum access time slot structure of the cognitive vehicle network model, a spectrum access strategy optimization model is constructed. The model is then solved using a continuous long short-term memory dual-duel deep Q-network to optimize the spectrum access strategy, reduce energy consumption, and improve the success rate.
It effectively reduces energy consumption, improves the success rate of spectrum access and communication efficiency, has good adaptability, and realizes efficient communication in cognitive vehicle networking.
Smart Images

Figure CN120018146B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of Internet of Vehicles communication, and in particular to a dynamic spectrum access method and device, equipment and medium. BACKGROUND
[0002] With the rapid development of intelligent transportation systems, Vehicular Ad-hoc Networks (VANETs) as the key to realize vehicle communication and automatic driving technology, has attracted widespread attention. Through wireless communication between vehicles, VANETs realizes information sharing and cooperation, and improves traffic efficiency and safety. However, VANETs need to run in a complex, changeable and highly dynamic wireless environment, which makes the scarcity and utilization efficiency of spectrum resources a problem to be solved.
[0003] Traditional static spectrum allocation methods cannot adapt to the rapidly changing communication needs of VANETs, resulting in low spectrum resource utilization and difficult to guarantee communication quality. Cognitive Radio (CR) technology provides a solution for efficient use of idle spectrum for VANETs through Dynamic Spectrum Access (DSA). Cognitive radio can intelligently perceive environmental changes and dynamically adjust spectrum usage strategies to achieve flexible allocation of spectrum resources, thereby improving communication efficiency. Specifically, cognitive radio can dynamically monitor spectrum usage, identify idle spectrum, and allocate these spectrum resources to secondary users (SUs) without interfering with primary users (PUs) through three functional modules of spectrum sensing, spectrum decision and spectrum access. However, how to effectively implement dynamic spectrum access in a highly dynamic VANET environment still faces many technical challenges.
[0004] First, frequent spectrum detection and switching operations will significantly increase device energy consumption, which not only increases the energy consumption of vehicles, but also puts higher requirements on their endurance. In VANETs, vehicles need to continuously monitor the surrounding spectrum environment and respond quickly, and this frequent operation will significantly increase the energy consumption of the device. Therefore, how to effectively reduce energy consumption while ensuring spectrum access efficiency is a problem to be solved.
[0005] Secondly, due to the dynamic changes and fierce competition of spectrum resources, vehicles often face access failure problems when trying to access the spectrum. This situation not only affects the communication quality of the Internet of Vehicles, but also can cause transmission delay or loss of critical information. For example, when a traffic accident or emergency occurs, vehicles need to communicate quickly to coordinate actions, but spectrum access failure can cause information to be unable to be delivered in time, thereby affecting traffic safety and efficiency. Therefore, improving the success rate of spectrum access is another important problem to be solved.
[0006] In addition, the high dynamics and complexity of cognitive Internet of Vehicles make the channel utilization rate low, making it difficult to guarantee the high throughput communication demand. The high-speed movement of vehicles and frequent changes in network topology make it difficult for traditional static spectrum access methods to adapt to the complex environment in actual applications. When on a highway or urban road, the moving speed and direction of vehicles change frequently, which puts higher requirements on the flexibility and adaptability of spectrum access strategies. SUMMARY
[0007] The purpose of the present application is to provide a dynamic spectrum access method, device, equipment and medium, which can effectively reduce energy consumption, improve the success rate of spectrum access, and realize efficient communication of cognitive Internet of Vehicles while ensuring spectrum access efficiency, and has good adaptability.
[0008] To achieve the above purpose, the present application provides the following solutions:
[0009] In a first aspect, the present application provides a dynamic spectrum access method, comprising:
[0010] determining a cognitive Internet of Vehicles network model in a two-lane scenario; the cognitive Internet of Vehicles network model comprises a roadside information unit and a plurality of vehicle users; the communication between the vehicle users and the roadside information unit is primary user communication; the communication between the vehicle users is secondary user communication;
[0011] determining a dynamic spectrum access time slot structure of the cognitive Internet of Vehicles network model; each time slot in the dynamic spectrum access time slot structure comprises a receiving stage and an access stage; the receiving stage is a stage in which the secondary user communication link receives the spectrum information sent by the roadside information unit; the access stage is a stage in which the secondary user communication link selects spectrum for access according to the received spectrum information;
[0012] constructing a spectrum access strategy optimization model based on the dynamic spectrum access time slot structure; the spectrum access strategy optimization model comprises an objective function and a constraint condition; the objective function is constructed with the minimum energy consumption and throughput of the secondary user communication link as the target; the constraint condition is determined according to the interference threshold, the energy consumption of the secondary user communication link and the spectrum access time of the secondary user communication link;
[0013] The spectrum access strategy optimization model is solved by using a long short-term memory double duel deep Q network to obtain an optimal spectrum access strategy of the secondary user communication link.
[0014] In a second aspect, the present application provides a dynamic spectrum access device, comprising:
[0015] A network model determination module is configured to determine a cognitive vehicle networking network model in a two-lane scenario, wherein the cognitive vehicle networking network model comprises roadside information units and a plurality of vehicle users, the communication between the vehicle users and the roadside information units is primary user communication, and the communication between the vehicle users is secondary user communication.
[0016] A time slot structure determination module is configured to determine a dynamic spectrum access time slot structure of the cognitive vehicle networking network model, wherein each time slot in the dynamic spectrum access time slot structure comprises a receiving stage and an access stage, the receiving stage is a stage in which a secondary user communication link receives spectrum information sent by a roadside information unit, and the access stage is a stage in which the secondary user communication link selects spectrum for access according to the received spectrum information.
[0017] An access strategy optimization model construction module is configured to construct a spectrum access strategy optimization model based on the dynamic spectrum access time slot structure, wherein the spectrum access strategy optimization model comprises a target function and a constraint condition, the target function is constructed with the minimum energy consumption and throughput of a secondary user communication link as a target, and the constraint condition is determined according to an interference threshold, the energy consumption of the secondary user communication link, and the spectrum access time of the secondary user communication link.
[0018] An optimization solving module is configured to solve the spectrum access strategy optimization model by using a long short-term memory double duel deep Q network to obtain an optimal spectrum access strategy of the secondary user communication link, wherein the long short-term memory double duel deep Q network comprises a long short-term memory network and a double duel deep Q network connected in sequence, and the double duel deep Q network is constructed based on a double deep Q network and a duel deep Q network with continuous learning introduced.
[0019] In a third aspect, the present application provides a computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the dynamic spectrum access method described above.
[0020] In a fourth aspect, the present application provides a computer readable storage medium, which stores a computer program. The computer program is executed by a processor to implement the dynamic spectrum access method.
[0021] According to the embodiments provided in the present application, the following technical effects are achieved.
[0022] The present application provides a dynamic spectrum access method, device, equipment and medium. In the dynamic spectrum access time slot structure of the cognitive vehicle networking network model, each time slot includes a receiving stage and an access stage. The secondary user communication link receives the spectrum information sent by the roadside information unit to select the spectrum for access. The characteristics of the roadside information unit being fixed and not consuming vehicle energy are utilized. The problem of different vehicles having different spectrum sensing results is improved. The energy consumption of spectrum sensing is effectively alleviated. The energy consumption and throughput of the secondary user communication link are minimized. The constraint conditions determined according to the interference threshold, the energy consumption of the secondary user communication link and the spectrum access time are used as constraints to build a spectrum access strategy optimization model. The throughput is maximized on the premise of meeting the energy consumption and interference constraints. Efficient communication in the cognitive vehicle networking is achieved. A continuous long short-term memory double duel deep Q network is used to solve the spectrum access strategy optimization model. The optimal spectrum access strategy of the secondary user communication link is obtained. The success rate of spectrum access is improved. The problems of slow model updating and poor adaptability are effectively overcome. BRIEF DESCRIPTION OF DRAWINGS
[0023] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application. Those skilled in the art can obtain other drawings according to these drawings without creative labor.
[0024] Figure 1 FIG. 1 is a diagram of an application environment of a dynamic spectrum access method according to an embodiment of the present application;
[0025] Figure 2 FIG. 2 is a flowchart of a dynamic spectrum access method according to an embodiment of the present application;
[0026] Figure 3 FIG. 3 is a schematic diagram of a cognitive vehicle networking network model according to an embodiment of the present application;
[0027] Figure 4 FIG. 4 is a schematic diagram of a dynamic spectrum access time slot structure according to an embodiment of the present application;
[0028] Figure 5A model structure schematic diagram of the D3QN network provided by an embodiment of the present application is shown in FIG. 1.
[0029] Figure 6 A structure schematic diagram of the LSTM network provided by an embodiment of the present application is shown in FIG. 2.
[0030] Figure 7 A schematic diagram of the CLD3QN-based DSA algorithm provided by an embodiment of the present application is shown in FIG. 3.
[0031] Figure 8 A functional module schematic diagram of a dynamic spectrum access apparatus provided by another embodiment of the present application is shown in FIG. 4.
[0032] Figure 9 A structure schematic diagram of a computer device provided by an embodiment of the present application is shown in FIG. 5. DETAILED DESCRIPTION
[0033] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.
[0034] The above-mentioned purposes, features and advantages of the present application can be more obvious and easy to understand. The present application will be described in further detail below with reference to the drawings and specific embodiments.
[0035] The dynamic spectrum access method provided by the embodiments of the present application can be applied to, for example, Figure 1The application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store the data required by the server 104 to process. The data storage system can be set up separately, or integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the cognitive vehicle networking network model to the server 104, and the server 104 receives the cognitive vehicle networking network model. For the cognitive vehicle networking network model, the server 104 determines the dynamic spectrum access time slot structure of the cognitive vehicle networking network model; based on the dynamic spectrum access time slot structure, a spectrum access strategy optimization model is constructed; the continuous long short-term memory double duel deep Q network is used to solve the spectrum access strategy optimization model, and the optimal spectrum access strategy of the secondary user communication link is obtained. The server 104 can feed back the optimal spectrum access strategy of the secondary user communication link to the terminal 102. In addition, in some embodiments, the dynamic spectrum access method can also be implemented by the server 104 or the terminal 102 alone, such as can be directly processed by the terminal 102 for the cognitive vehicle networking network model, or the server 104 can obtain the cognitive vehicle networking network model from the data storage system and process it for the cognitive vehicle networking network model.
[0036] Among them, the terminal 102 can be, but not limited to, various desktop computers, notebook computers, smart phones, tablet computers, Internet of Things devices and portable wearable devices. The Internet of Things device can be a smart speaker, a smart television, a smart air conditioner, a smart vehicle device, etc. The portable wearable device can be a smart watch, a smart bracelet, a head-mounted device, etc. The server 104 can be implemented by an independent server or a server cluster composed of multiple servers, and can also be a cloud server.
[0037] In an exemplary embodiment, as Figure 2 shown, a dynamic spectrum access method is provided, which is executed by a computer device, specifically by a terminal or a server, etc. Computer device alone, or by a terminal and a server together, in the embodiment of the present application, take the server 104 in Figure 1 for example, the following steps 201 to 204 are included. Among them:
[0038] Step 201, determine the cognitive vehicle networking network model in the double lane scene.
[0039] Among them, the cognitive vehicle networking network model includes: a road side information unit (Road Side Unit, RSU) and a plurality of vehicle users; the communication between the vehicle user and the road side information unit is the primary user communication; the communication between the vehicle users is the secondary user communication.
[0040] In step 202, a dynamic spectrum access time slot structure of the cognitive vehicle networking network model is determined.
[0041] Each time slot in the dynamic spectrum access time slot structure includes a receiving stage and an access stage; the receiving stage is a stage in which the secondary user communication link receives spectrum information sent by the roadside information unit; and the access stage is a stage in which the secondary user communication link selects spectrum for access according to the received spectrum information.
[0042] In step 203, a spectrum access strategy optimization model is constructed based on the dynamic spectrum access time slot structure.
[0043] The spectrum access strategy optimization model includes a target function and a constraint condition; the target function is constructed with the minimum energy consumption and throughput of the secondary user communication link as the target; and the constraint condition is determined according to the interference threshold, the energy consumption of the secondary user communication link, and the spectrum access time of the secondary user communication link.
[0044] In step 204, a continual LSTM double dueling deep Q-network (CLD3QN) is used to solve the spectrum access strategy optimization model, so as to obtain an optimal spectrum access strategy of the secondary user communication link.
[0045] The continual LSTM double dueling deep Q-network includes a long short-term memory network and a double dueling deep Q-network connected in sequence; the double dueling deep Q-network is constructed by introducing continual learning (CL) and based on a double deep Q-network (DDQN) and a dueling deep Q-network (DQN).
[0046] Implementing the above steps 201 to 204 can effectively reduce energy consumption, improve the success rate of spectrum access, realize efficient communication of the cognitive vehicle network, and adapt well while ensuring spectrum access efficiency.
[0047] In another exemplary embodiment of the present application, the cognitive vehicle networking network model in step 201 and the dynamic spectrum access time slot structure of the cognitive vehicle networking network model in a double-lane scenario are further introduced.
[0048] Considering the complexity of future networks, a cognitive vehicle networking network of multiple vehicle users and roadside information units is designed in an urban environment, and the cognitive vehicle networking network model is as follows Figure 3As shown, this embodiment implements vehicle-to-everything (V2X) communication in a two-lane urban environment, including a base station (BS), roadside information units (RSUs), and vehicle users. The main communication methods are vehicle-to-integrated (V2I) communication between RSUs and vehicle users, and vehicle-to-vehicle (V2V) communication between vehicle users. In the model, V2V is defined as secondary user communication, and V2I is defined as primary user communication. To facilitate efficient communication and collision-free transmission, each primary user communication is allocated a separate radio channel. Figure 3 In this context, T represents a time slot, R represents reception, and A represents access.
[0049] Assume the licensed frequency band set for V2I communication is l∈{1,2,…,L}, and the set for V2V communication is n∈{1,2,…,N}, where L and N represent the number of V2I and V2V communications, respectively. Furthermore, assume the current amount of spectrum in the vehicular network environment is the same as the licensed spectrum, and the channel combination in the current system model is m∈{1,2,…,M}, with negligible cross-channel interference. To improve spectrum resource utilization, secondary user communication opportunities will be provided to access idle frequency bands licensed but not used by the primary user.
[0050] like Figure 4 As shown in section (a), in the traditional time-slot structure of Dynamic Spectrum Access (DSA), the V2V link needs to sense spectrum gaps on its own before selecting to access idle spectrum. This time-slot structure makes it difficult to fully meet the communication sensing requirements of different vehicles and wastes too much energy of the V2V link for frequent spectrum sensing. This embodiment utilizes the fixed nature of roadside information units (BITs) that do not consume vehicle energy, improving the problem of different vehicles having different spectrum sensing results and effectively alleviating the energy consumption of spectrum sensing.
[0051] Therefore, this embodiment proposes the following... Figure 4 The improved DSA timeslot structure shown in section (b) allows V2V links to receive spectrum information from roadside information units before accessing the spectrum, and then select a suitable spectrum for access. During this process, the access time τ of each V2V link is limited to the current timeslot, satisfying the constraint τ≥0. Furthermore, to ensure the rational use of spectrum resources, the access time τ should be less than the total timeslot T, satisfying the constraint T-τ≥0, where T represents the duration of the entire timeslot. This design ensures that each V2V link can complete the access operation within a limited time and avoids communication conflicts or resource waste due to exceeding the timeslot time.
[0052] In another exemplary embodiment of this application, the idea of constructing the spectrum access strategy optimization model in step 203 is introduced.
[0053] According to the improved dynamic spectrum access time slot structure, the energy consumption formula of the secondary user communication link n on the channel m can be expressed as:
[0054]
[0055] Wherein, τ represents the spectrum access time of the secondary user communication link, that is, the time consumed by access, P n represents the transmission power of the secondary user communication link n, represents the channel gain of the secondary user communication link n on the channel m. At the same time, in order to optimize the energy consumption, the secondary user communication link also needs to meet the constraint condition of maximum energy consumption in the access process, so as to ensure that the system controls the energy consumption while maximizing the throughput, which is expressed as
[0056] In the considered vehicle networking environment, it is assumed that the propagation condition between the communication users is line of sight (LOS) environment, and the channel model is Rician fading channel. Therefore, the channel gain of the secondary user communication link n on the channel m The calculation formula is as follows:
[0057]
[0058] Wherein, k represents the ratio of received signal power to scattered path power variance under LOS, and θ represents the phase of received signal. CN(·) is a circular symmetric complex Gaussian random variable, is the scale parameter of the secondary user communication link n. The path loss of the secondary user communication link n is related to the communication distance distance and the carrier frequency f c The calculation formula of the path loss of the secondary user communication link n is as follows:
[0059]
[0060] Wherein, represents the communication distance of the secondary user communication link n, f c represents the carrier frequency of the wireless channel, PL α , PL β and PL γ respectively represent the path loss of the reference distance, the path loss index and the frequency correlation of the path loss, GHz represents the unit of the carrier frequency of the line channel.
[0061] In the model of the embodiment, it is assumed that the transceivers of the RSU and the vehicle both use one antenna, and respectively, represent the position information of RSUi and vehicle j, where i e {1, 2,..., I} and j e {1, 2,..., J}. Therefore, the communication distance between secondary user communication link n and another vehicle j' and vehicle j is calculated as The formula is as follows:
[0062]
[0063] represents the position information of vehicle j'.
[0064] Meanwhile, the channel gain of primary user communication link l as a primary user communication on channel m is calculated as follows:
[0065]
[0066] wherein is the scale parameter of primary user communication link l. The path loss of primary user communication link l is related to the communication distance and the carrier frequency f c , and the formula is as follows:
[0067]
[0068] wherein represents the communication distance of primary user communication link l. The calculation formula of the communication distance between roadside information unit i and vehicle j of primary user communication link l is as follows:
[0069]
[0070] After considering the improved dynamic spectrum access time slot structure and its energy consumption constraints, the embodiment next focuses on the impact of interference on the access of secondary user communication link. The throughput of secondary user communication link n when accessing different channels m can be represented as:
[0071]
[0072] wherein represents the current access strategy of secondary user communication link to channel m, B m represents the bandwidth of channel m, represents the transmission power of secondary user communication link n, and will change in size according to the access strategy , represents the noise power spectral density. According to formula (8), the access strategy has a significant impact on the throughput.
[0073] Therefore, the embodiment designs a new hybrid access mode, which combines the characteristics of Overlay and Underlay two different access modes. After receiving the channel information sent by the RSU, the SU makes a decision to access the channel m according to the channel state information. If the sensing error information appears at this time, the SU judges whether the current channel exceeds the interference threshold I th of the current channel, and if not, the transmission power is reduced to complete the access process, otherwise the transmission fails. The spectrum access strategy of the secondary user communication link accessing the channel m under different interference, i.e., the secondary user communication link accessing the channel m is as follows:
[0074]
[0075] wherein, I represents the interference of other primary user communication links l to the secondary user communication link n on the channel m at the current time t, and the calculation formula is as follows:
[0076]
[0077] respectively represent the interference threshold when there is one primary user communication link, the interference threshold when there is one secondary user communication link, and the maximum interference threshold. When there is no interference on the current channel m, the channel can be successfully accessed. When there is a primary user communication link on the channel m and the interference is less than the interference threshold I of the primary user communication link, the secondary user communication link can successfully access the current channel and does not need to reduce the power; otherwise, the power needs to be reduced to ensure successful access of the channel. When there is a secondary user communication link on the channel m and the interference is less than the interference threshold I of the secondary user communication link, the secondary user communication link can successfully access the current channel and does not need to reduce the power; otherwise, the power needs to be reduced to ensure successful access of the channel. When there is an interference greater than the maximum interference threshold I on the channel m, the access fails.
[0078] Therefore, the optimization problem can represent the secondary user communication link n to find the best access strategy in all licensed channels m, to maximize the throughput of the secondary user communication link and the lowest energy consumption, and to jointly optimize the access time, the interference threshold, and to construct the spectrum access strategy optimization model.
[0079] Based on the above, step 203 specifically includes:
[0080] (1) determining the number of spectrums based on the dynamic spectrum access time slot structure; each spectrum corresponds to a channel.
[0081] (2) Calculate the communication distance of two vehicle users in the secondary user communication link, the calculation formula is shown in formula (4).
[0082] (3) Calculate the path loss of the secondary user communication link according to the communication distance of two vehicle users in the secondary user communication link and the carrier frequency, the calculation formula is shown in formula (3).
[0083] (4) Calculate the channel gain of the secondary user communication link on each channel according to the path loss of the secondary user communication link, the calculation formula is shown in formula (2).
[0084] (5) Calculate the energy consumption of the secondary user communication link on each channel according to the channel gain of the secondary user communication link on each channel, the transmission power of the secondary user communication link and the spectrum access time of the secondary user communication link, the calculation formula is shown in formula (1).
[0085] (6) Calculate the communication distance between the vehicle user and the roadside information unit in the primary user communication link, the calculation formula is shown in formula (7).
[0086] (7) Calculate the path loss of the primary user communication link according to the communication distance between the vehicle user and the roadside information unit in the primary user communication link and the carrier frequency, the calculation formula is shown in formula (6).
[0087] (8) Calculate the channel gain of the primary user communication link on each channel according to the path loss of the primary user communication link, the calculation formula is shown in formula (5).
[0088] (9) Calculate the interference of the primary user communication link to the secondary user communication link on each channel according to the channel gain of the primary user communication link on each channel, the calculation formula is shown in formula (10).
[0089] (10) Determine the spectrum access strategy of the secondary user communication link to access each channel according to the interference of the primary user communication link to the secondary user communication link and the interference threshold, the calculation formula is shown in formula (9).
[0090] (11) Calculate the throughput of the secondary user communication link when accessing each channel according to the spectrum access strategy of the secondary user communication link to access each channel, the bandwidth of the channel, the spectrum access time of the secondary user communication link, the channel gain of the primary user communication link on each channel, the transmission power of the secondary user communication link and the noise power spectral density, the calculation formula of the throughput is shown in formula (8).
[0091] (12) Build a spectrum access strategy optimization model based on the energy consumption of the secondary user communication link on each channel, the throughput of the secondary user communication link when accessing each channel, the interference threshold and the spectrum access time of the secondary user communication link.
[0092] The expression of the spectrum access strategy optimization model is:
[0093]
[0094] Wherein, max represents taking the maximum value; s.t. represents the constraint condition; N represents the number of secondary user communication links; M represents the number of channels; represents the throughput of the secondary user communication link n when accessing the channel m; represents the spectrum access strategy of the secondary user communication link accessing the channel m; I th represents the interference threshold; τ represents the spectrum access time of the secondary user communication link; represents the energy consumption of the secondary user communication link n on the channel m; represents the interference threshold of the secondary user communication link; represents the interference threshold of the primary user communication link; represents the maximum interference threshold; T represents the entire time for completing the access process. In this embodiment, characterizes different strategies of the current secondary user communication link accessing the channel m, and the energy consumption of the secondary user communication link in the designed hybrid access mode has a maximum value constraint, an interference threshold has a maximum minimum value constraint.
[0095] Formula (11) is a nonlinear optimization problem containing multiple constraints, and it is difficult to solve directly. Due to the dynamic changes of the Internet of Vehicles environment, vehicles face different interference thresholds and energy consumption constraints when making DSA decisions, and it is difficult to obtain the optimal spectrum access strategy in real time. In addition, in a high-density traffic scenario, the competition for limited spectrum resources between vehicles may cause serious interference and reduce the overall transmission efficiency. Therefore, an enhanced dynamic spectrum access algorithm (step 204) is designed to maximize the throughput under the premise of meeting the energy consumption and interference constraints, and to realize efficient communication in the cognitive Internet of Vehicles.
[0096] In another exemplary embodiment of the application, the dynamic spectrum access algorithm based on CLD3QN in step 204 is introduced. For convenience, the secondary user communication is represented by V2V, and the primary user communication is represented by V2I.
[0097] In the system model of Figure 3 , multiple V2V links try to access the limited spectrum occupied by V2I links, which can be modeled as a multi-agent RL problem. Each V2V link acts as an agent, interacts with the unknown communication environment to obtain experience, and then uses the experience to guide its own strategy design. Multiple V2V agents jointly explore the environment and improve the spectrum access strategy according to their own observations of the environment state.
[0098] In this embodiment, CLD3QN is used to realize distributed DSA of V2V link in dynamic vehicle environment. LSTM is mainly used to save historical spectrum decision data of V2V, D3QN is mainly used to determine the optimal spectrum access strategy of V2V, and CL is used to adapt to new changes. V2V link first obtains channel state from RSU, selects available channel for data transmission, and obtains corresponding reward. Then, the channel state, the action at the current time and the reward are input into the LSTM network. Finally, the calculation result of the LSTM network is input into the network of D3QN to estimate the optimal channel selection strategy.
[0099] 1) D3QN network.
[0100] D3QN is an improvement of double deep Q network (DDQN) and duel deep Q network (Dueling DQN). D3QN decomposes the Q value function into state value (StateValue) and advantage (Advantage) two parts, so that the estimation is more accurate, and the learning effect is improved. The model structure of D3QN network is shown in Figure 5 Dynamic spectrum access in cognitive Internet of vehicles environment involves time-varying and spatial variability of spectrum resources. This method is particularly suitable for complex environments of dynamic spectrum access, and can effectively improve spectrum utilization.
[0101] D3QN uses two deep neural networks for data learning, namely evaluation Q-network (EQN) and target Q-network (TQN). The evaluation Q-network obtains the current state and action samples from the experience pool, and predicts the Q value of the specific action. The target Q-network obtains the next state sample from the experience pool, and predicts the best Q value from all actions starting from this state. The evaluation Q-network and the target Q-network are the same in network structure, and are divided into state value function and advantage function when solving Q value function, so as to better estimate Q value.
[0102] In D3QN, the evaluation Q-network and the target Q-network both use duel network structure, and decompose the Q value function into state value function V(s,ω) and advantage function A(s,a,ω). Wherein, s is the current state, a is the action, and ω is the network weight. The Q value function can be expressed as:
[0103]
[0104] Where |A| is the size of the action set, ω is the parameter of the evaluation Q-network, and α is the parameter for balancing the state value function and the action advantage function.
[0105] The target Q value update formula in D3QN is:
[0106]
[0107] where ω' is the parameter of the target Q-network.
[0108] The Q-value update policy in D3QN can be described as:
[0109]
[0110] where r is the reward of the action selected for the current state, γ∈(0,1) is the Q-value discount rate of D3QN, and β∈(0,1) is the learning rate.
[0111] The selected action is determined by an ε-greedy policy, as follows:
[0112]
[0113] The D3QN algorithm uses an experience replay mechanism to store the experience transition samples obtained by the agent from the environment at each time. After a number of steps, a batch of samples is randomly drawn from the experience pool. Then, D3QN uses the mini-batch stochastic gradient descent method to update the network parameters ω to minimize the loss function L(ω), which is expressed as follows:
[0114]
[0115] During the training process, the weights of EQN are updated through gradient descent and backpropagation techniques. The target network TQN remains unchanged. After a number of training rounds, the weights of EQN are copied to TQN. This helps TQN to obtain improved weights, so that it can also predict more accurate target Q-values.
[0116] 2) LSTM network.
[0117] In order to learn an efficient spectrum access strategy based on historical experience, the embodiment uses an LSTM network to save long-term historical spectrum selection results and solve the problems of gradient vanishing and gradient explosion in long sequence training process. In the process of reinforcement learning, a large amount of historical information can help the vehicle quickly understand the characteristics of the cognitive vehicle networking environment. As shown in Figure 6 The LSTM network consists of three gates. The first is the forget gate, which determines which information in the memory cell needs to be forgotten. Its calculation formula is:
[0118] f t =σ(W f ·[h t-1 ,x t ]+b f )(22)
[0119] where f t is the output of the forget gate and f t ∈(0, 1), σ is the Sigmoid function, W f is the weight matrix of the forget gate, h t-1 is the hidden state at the previous time step, x t is the input at the current time step, and b f is the bias term of the forget gate. The second is the input gate, which determines how much new information is added to the memory cell, and its calculation formula is:
[0120] i t = σ(W i · [h t-1 , x t ] + b i )(23)
[0121] where i t is the output of the input gate, W i is the weight matrix of the input gate, and b i is the bias term of the input gate; the candidate value of new information is generated through a tanh layer:
[0122] C' t = tanh(W C · [h t-1 , x t ] + b C )(24)
[0123] where C' t is the new candidate memory, tanh is the hyperbolic tangent function, W C represents the weight matrix of the candidate memory, and b C is the bias term of the candidate memory; after obtaining the candidate value, the memory cell (Cell State) is updated, and the state of the memory cell is updated by combining the forget gate and the input gate:
[0124] C t = f t · C t-1 + i t · C' t (25)
[0125] where C t is the memory cell state at the current time step, and C t-1 is the memory cell state at the previous time step. The third is the output gate (Output Gate), which determines the hidden state (Hidden State) at the current time step, and its calculation formula is:
[0126] o t= σ(W o · [h t-1 , x t ] + b o )(26)
[0127] where o t is the output of the output gate, W o is the weight matrix of the output gate, and b o is the bias term of the output gate; the hidden state is obtained by combining the memory cell state and the output gate:
[0128] h t = o t · tanh(C t )(27)
[0129] where h t represents the hidden state at the current time.
[0130] In summary, the input x t and the hidden state h t-1 and the memory cell state C t-1 at the previous time are calculated through the above formula, and the hidden state h t and the memory cell state C t at the current time are output.
[0131] 3) Dynamic spectrum access (DSA) algorithm based on CLD3QN.
[0132] In order to realize the intelligent spectrum access of cognitive vehicle networking, this embodiment proposes a dynamic spectrum access algorithm for cognitive vehicle networking based on CLD3QN. LSTM network is good at processing time series data and can capture the historical information and time dependence of spectrum use, while CL can adapt to dynamic environment more efficiently. By using two Q networks (evaluation Q network and target Q network) to separate action selection and Q value update, it can effectively reduce the overestimation problem of Q value and improve the stability of learning. In the two Q networks, the Q value function is decomposed into state value function and advantage function, which can more accurately estimate the value of the state and the relative advantage of different actions. This decomposition method can reduce the variance of Q value estimation, improve the stability and accuracy of decision-making. By combining CL and LSTM, the model can adapt to changes in dynamic spectrum environment and continuously update and improve the spectrum access strategy. After the idea of continuous learning is added, the new loss function changes to:
[0133]
[0134] where L CLD3QN (ω') represents the loss function; represents the expected value of sampling data from the experience replay pool; Q e(s, a; ω) represents the Q value evaluated by the Q network output; Q t (s, a) represents the Q value output by the target Q network; ψ || ω'- ω || represents the difference between the current duel deep Q network parameter and the last batch evaluation Q network parameter 2 represents L2 regularization (also known as weight decay) for preventing model overfitting; ψ represents the regularization parameter; s represents the channel state; a represents the action; r represents the reward; s' represents the next channel state; ω represents the last batch evaluation Q network parameter, which is used to constrain the change between the new and old parameters to avoid catastrophic forgetting; ω' represents the current duel deep Q network parameter.
[0135] Based on Figure 4 Based on the time slot structure shown in FIG. 1, a CLD3QN-based overlay spectrum access scheme is proposed. In DSA, V2V links can only use the idle channels that are not occupied by V2I links. If a channel is being used by V2I, V2I cannot access the channel in any way. In time slot T, vehicles obtain the state of M channels through the broadcast of RSU. Let binary variable represent the state of channel m, then the channel state obtained by V2V link can be represented by vector , where represents that the channel is in a busy state in the current period, represents that the channel is in an idle state in the current period, and represents the interference threshold of the current channel. It is assumed that the vehicle selects only one channel for communication in each time slot. After obtaining the channel state, the action taken by V2V link is to access a certain channel, where represents that V2V link selects to access channel m in the current period, represents that V2V link does not select to access channel m in the current period.
[0136] When V2V link successfully enters a channel and completes transmission in a time slot, it will obtain a learning reward. For V2V link n in period t, the reward R t in state S t and action A t is represented as:
[0137]
[0138] where Γ m , represent the throughput, energy consumption and interference of V2V link n in the current channel m, χ is the energy consumption coefficient, and ι is the interference coefficient.
[0139] Therefore, under different access conditions, the reward function of V2V link can be represented as:
[0140]
[0141] wherein R t (S t , A t ) represents the channel state at current time t and the reward under action at current time t; represents the throughput of secondary user communication link n when accessing channel m; represents the energy consumption of secondary user communication link n on channel m; χ represents the energy consumption coefficient; ι represents the interference coefficient; represents the interference of primary user communication link on secondary user communication link n on channel m; represents the interference threshold of primary user communication link; represents the interference threshold of secondary user communication link; C represents a constant; represents the interference of primary user communication link on secondary user communication link n on channel m at current time t.
[0142] The DSA algorithm based on CLD3QN proposed in the embodiment is as shown in Figure 7 The LSTM network stores the long-term historical observation data perceived by the RSU, and the D3QN calculates the optimal spectrum access strategy of the V2V link, and updates the continuous learning every 10 steps to adapt to the dynamic spectrum environment. In order to facilitate understanding, the embodiment illustrates a simple DSA process in the algorithm process. As shown in Figure 7 , the RSU perceives the channel state from the environment, and broadcasts the channel state to all vehicle users, V2VAgent1 selects to access channel 2 according to the received information. The perceived state of channel 2 is “0”, indicating that channel 2 is available; the actual state is “1”, indicating that channel 3 is occupied by V2I.
[0143] Therefore, V2VAgent2 collides with V2I after accessing channel 3, and if the interference does not exceed the reward is otherwise the reward is V2VAgent2 selects to access channel 3 according to the received information. The perceived state of channel 3 is “0”, indicating that channel 3 is available; the actual state is “0”, indicating that channel 3 is not occupied by V2I. Next, V2VAgent2 successfully accesses channel 3, and the reward is At the same time, V2VAgentN also selects to access channel 3 according to the received information. As known above, V2VAgentN also selects channel 3 to access, so V2VAgent3 collides with V2VAgent1 after accessing channel 2, and if the reward is otherwise the reward is When more than two users, i.e. a third user, attempt to access the channel, the reward is defined as -C, where C is a constant, which means that the communication quality of all three users will be severely interfered and broken. The reward is 0 when the agent does not select a channel.
[0144] The pseudo code of the CLD3QN-based DSA algorithm is described as follows:
[0145] Algorithm 1: CLD3QN-based DSA algorithm
[0146] Input: Initialize the CLD3QN network of each SU, set the learning rate β, the discount factor γ, the action selection policy and the false alarm probability Pf, and initialize the parameters α of the continual learning model.
[0147] Output: The optimal spectrum access strategy of each SU.
[0148]
[0149] Based on the above description, step 204 specifically includes:
[0150] (1) Obtain the channel state perceived by the roadside information unit from the environment.
[0151] (2) Based on the spectrum access strategy optimization model, when the secondary user communication link selects a channel according to the received channel state and completes data transmission in a time slot, calculate the reward under the channel state at the current time and the action at the current time. The calculation formula of the reward under the channel state at the current time and the action at the current time is shown in formula (30).
[0152] (3) Input the channel state at the current time, the action at the current time and the reward into the long short-term memory network to obtain the channel state at the next time.
[0153] (4) Construct a double duel deep Q network; the double duel deep Q network includes an evaluation Q network and a target Q network.
[0154] (5) Input the channel state at the current time and the action at the current time into the evaluation Q network, input the channel state at the next time into the target Q network, train the double duel deep Q network with the loss function as the target, and output the Q value corresponding to the spectrum access strategy of the target Q network as the optimal spectrum access strategy of the secondary user communication link; the loss function is determined by introducing continual learning and according to the reward, the Q value output by the evaluation Q network and the Q value output by the target Q network. The expression of the loss function is shown in formula (28).
[0155] The dynamic spectrum access method of the present application is a green and efficient dynamic spectrum access method realized based on a continuous long short-term memory double duel deep Q network (CLD3QN), successfully combines CR with VANETs, and significantly improves the utilization rate of spectrum resources. By introducing a hybrid access mode, the Overlay and Underlay technologies are comprehensively utilized, so that the access success rate of the system in different user number scenarios is optimized. The proposed CLD3QN algorithm combines continuous learning and long short-term memory network, and effectively overcomes the problems of slow model updating and poor adaptability in traditional methods.
[0156] Based on the same inventive concept, the embodiments of the present application also provide a dynamic spectrum access device for implementing the dynamic spectrum access method described above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme described in the above method, so the specific limitations in one or more dynamic spectrum access device embodiments provided below can refer to the limitations of the dynamic spectrum access method described above, which will not be repeated here.
[0157] In one exemplary embodiment, as shown in Figure 8 A dynamic spectrum access device is provided, comprising:
[0158] The network model determination module 801 is configured to determine a cognitive vehicle networking network model in a two-lane scenario. The cognitive vehicle networking network model comprises a roadside information unit and a plurality of vehicle users. The communication between the vehicle users and the roadside information unit is primary user communication, and the communication between the vehicle users is secondary user communication.
[0159] The time slot structure determination module 802 is configured to determine a dynamic spectrum access time slot structure of the cognitive vehicle networking network model. Each time slot in the dynamic spectrum access time slot structure comprises a receiving stage and an access stage. The receiving stage is a stage in which a secondary user communication link receives spectrum information sent by the roadside information unit. The access stage is a stage in which the secondary user communication link selects spectrum for access according to the received spectrum information.
[0160] The access strategy optimization model construction module 803 is configured to construct a spectrum access strategy optimization model based on the dynamic spectrum access time slot structure. The spectrum access strategy optimization model comprises an objective function and a constraint condition. The objective function is constructed with the minimum energy consumption and throughput of the secondary user communication link as the target. The constraint condition is determined according to the interference threshold, the energy consumption of the secondary user communication link, and the spectrum access time of the secondary user communication link.
[0161] The optimization solving module 804 is configured to solve the spectrum access strategy optimization model by using a long short-term memory double dueling deep Q network to obtain an optimal spectrum access strategy of the secondary user communication link. The long short-term memory double dueling deep Q network comprises a long short-term memory network and a double dueling deep Q network connected in sequence. The double dueling deep Q network is constructed based on a double deep Q network and a dueling deep Q network and introduces continuous learning.
[0162] In an exemplary embodiment, a computer device, which can be a server or a terminal, is provided. An internal structure diagram of the computer device can be as shown in Figure 9 The computer device includes a processor, a memory, an input / output interface (I / O), and a communication interface. The processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. The processor of the computer device is configured to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for running the operating system and the computer program in the non-volatile storage medium. The database of the computer device is configured to store a cognitive vehicle networking network model in a two-lane scenario. The input / output interface of the computer device is configured to exchange information between the processor and external devices. The communication interface of the computer device is configured to communicate with external terminals through network connection. The computer program is executed by the processor to implement a dynamic spectrum access method.
[0163] Those skilled in the art can understand that Figure 9 the structure shown in the above description is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the computer device to which the scheme of the present application is applied. Specifically, the computer device can include more or fewer components than those shown in the diagram, or combine certain components, or have a different arrangement of components. In an exemplary embodiment, a computer device is provided, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0164] In an exemplary embodiment, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.
[0165] In an exemplary embodiment, a computer program product is provided, which includes a computer program. The computer program is executed by a processor to implement the steps in the above method embodiments.
[0166] It should be noted that the user information (including but not limited to user equipment information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or authorized by all parties, and the collection, use and processing of related data need to comply with relevant regulations.
[0167] It can be understood by those skilled in the art that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing related hardware through a computer program. The computer program can be stored in a non-volatile computer readable storage medium, and when the computer program is executed, it can include the processes of the above-mentioned embodiments of each method. Any reference to memory, database or other medium used in each embodiment provided by the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (Read-Only Memory, ROM), magnetic tape, floppy disk, flash memory, optical storage, high-density embedded non-volatile memory, resistive memory (ReRAM), magnetoresistive random access memory (Magnetoresistive Random Access Memory, MRAM), ferroelectric memory (Ferroelectric Random Access Memory, FRAM), phase change memory (Phase Change Memory, PCM), graphene memory, etc. Volatile memory can include random access memory (Random Access Memory, RAM) or external cache memory, etc. As an illustration but not limitation, RAM can be in various forms, such as static random access memory (Static Random Access Memory, SRAM) or dynamic random access memory (Dynamic Random Access Memory, DRAM), etc.
[0168] The database involved in each embodiment provided by the present application can include at least one of a relational database and a non-relational database. The non-relational database can include a distributed database based on a block chain, etc., without being limited thereto. The processor involved in each embodiment provided by the present application can be a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic device, etc., without being limited thereto.
[0169] Each technical feature of the above embodiments can be combined arbitrarily. In order to make the description simple, not all possible combinations of each technical feature in the above embodiments are described, but as long as the combination of these technical features does not exist contradictory, it should be considered as the scope of the present disclosure.
[0170] The principles and implementation manners of the present application are described herein by using specific examples, and the above examples are only used to help understand the method of the present application and its core idea; meanwhile, for those skilled in the art, according to the idea of the present application, the specific implementation manners and application ranges will have changes. In conclusion, the content of the specification should not be understood as a limitation of the present application.
Claims
1. A dynamic spectrum access method, characterized in that, The dynamic spectrum access method includes: A cognitive vehicle-to-everything (V2X) network model is defined for a two-lane scenario. The cognitive V2X network model includes: a roadside information unit and multiple vehicle users; the communication between the vehicle users and the roadside information unit is primary user communication; the communication between the vehicle users is secondary user communication. The dynamic spectrum access time slot structure of the cognitive vehicle network model is determined; each time slot in the dynamic spectrum access time slot structure includes a receiving phase and an access phase; the receiving phase is the phase in which the secondary user communication link receives the spectrum information sent by the roadside information unit; the access phase is the phase in which the secondary user communication link selects a spectrum for access based on the received spectrum information. A spectrum access strategy optimization model is constructed based on the dynamic spectrum access time slot structure. The spectrum access strategy optimization model includes an objective function and constraints. The objective function is constructed with the goal of minimizing the energy consumption and throughput of the secondary user communication link. The constraints are determined based on the interference threshold, the energy consumption of the secondary user communication link, and the spectrum access time of the secondary user communication link. The spectrum access strategy optimization model is solved using a persistent long short-term memory dual-duel deep Q-network to obtain the optimal spectrum access strategy for the secondary user communication link. The persistent long short-term memory dual-duel deep Q-network includes a long short-term memory network and a dual-duel deep Q-network connected in sequence. The dual-duel deep Q-network is constructed by introducing continuous learning and based on the dual deep Q-network and the duel deep Q-network.
2. The dynamic spectrum access method according to claim 1, characterized in that, Based on the aforementioned dynamic spectrum access time slot structure, a spectrum access strategy optimization model is constructed, specifically including: The number of spectrum segments is determined based on the dynamic spectrum access time slot structure; each spectrum segment corresponds to one channel. Calculate the communication distance between two vehicle users in the secondary user communication link; The path loss of the secondary user communication link is calculated based on the communication distance between the two vehicle users and the carrier frequency in the secondary user communication link. Calculate the channel gain of the secondary user communication link on each channel based on the path loss of the secondary user communication link. Calculate the energy consumption of the secondary user communication link on each channel based on the channel gain, transmission power, and spectrum access time of the secondary user communication link on each channel. Calculate the communication distance between the vehicle user and the roadside information unit in the primary user communication link; The path loss of the primary user communication link is calculated based on the communication distance between the vehicle user and the roadside information unit and the carrier frequency in the primary user communication link. Calculate the channel gain of the primary user communication link on each channel based on the path loss of the primary user communication link; The interference of the primary user communication link to the secondary user communication link on each channel is calculated based on the channel gain of the primary user communication link on each channel. The spectrum access strategy for secondary user communication links to access each channel is determined based on the interference and interference threshold of the primary user communication link to the secondary user communication link. Based on the spectrum access strategy of the secondary user communication link to each channel, the bandwidth of the channel, the spectrum access time of the secondary user communication link, the channel gain of the primary user communication link on each channel, the transmit power and noise power spectral density of the secondary user communication link, calculate the throughput of the secondary user communication link when accessing each channel. A spectrum access strategy optimization model is constructed based on the energy consumption of the secondary user communication link on each channel, the throughput of the secondary user communication link when accessing each channel, the interference threshold, and the spectrum access time of the secondary user communication link.
3. The dynamic spectrum access method according to claim 1, characterized in that, The optimal spectrum access strategy for the secondary user communication link is obtained by solving the spectrum access strategy optimization model using a continuous long short-term memory dual-duel deep Q-network, specifically including: Acquire the channel state perceived by the roadside information unit from the environment; Based on the spectrum access strategy optimization model, when the secondary user communication link selects a channel according to the received channel state and completes data transmission within a time slot, the reward is calculated based on the channel state at the current moment and the action at the current moment. The current channel state, the current action, and the reward are input into the Long Short-Term Memory network to obtain the channel state at the next time step. Construct a dual-duel depth Q-network; the dual-duel depth Q-network includes: an evaluation Q-network and a target Q-network; The current channel state and the current action are input into the evaluation Q network, and the channel state at the next moment is input into the target Q network. The dual-duel depth Q network is trained with the goal of minimizing the loss function. The spectrum access policy corresponding to the Q value output by the trained target Q network is used as the optimal spectrum access policy for the secondary user communication link. The loss function is determined by introducing continuous learning and based on the reward, the Q value output by the evaluation Q network, and the Q value output by the target Q network.
4. A dynamic spectrum access device, characterized in that, The dynamic spectrum access device includes: A network model determination module is used to determine the cognitive vehicle network model in a two-lane scenario; the cognitive vehicle network model includes: a roadside information unit and multiple vehicle users; the communication between the vehicle users and the roadside information unit is primary user communication; the communication between the vehicle users is secondary user communication. A time slot structure determination module is used to determine the dynamic spectrum access time slot structure of the cognitive vehicle network model; each time slot in the dynamic spectrum access time slot structure includes a receiving phase and an access phase; the receiving phase is the phase in which the secondary user communication link receives spectrum information sent by the roadside information unit; the access phase is the phase in which the secondary user communication link selects spectrum for access based on the received spectrum information; An access strategy optimization model construction module is used to construct a spectrum access strategy optimization model based on the dynamic spectrum access time slot structure. The spectrum access strategy optimization model includes an objective function and constraints. The objective function is constructed with the goal of minimizing the energy consumption and throughput of the secondary user communication link. The constraints are determined based on the interference threshold, the energy consumption of the secondary user communication link, and the spectrum access time of the secondary user communication link. The optimization solution module is used to solve the spectrum access strategy optimization model using a persistent long short-term memory dual-duel deep Q-network to obtain the optimal spectrum access strategy for the secondary user communication link. The persistent long short-term memory dual-duel deep Q-network includes a long short-term memory network and a dual-duel deep Q-network connected in sequence. The dual-duel deep Q-network is constructed by introducing continuous learning and based on the dual deep Q-network and the duel deep Q-network.
5. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the dynamic spectrum access method according to any one of claims 1-3.
6. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the dynamic spectrum access method according to any one of claims 1-3.
Citation Information
Patent Citations
Dynamic spectrum access method based on bidirectional long-short-term memory network
CN112672359A
Cognitive radio network dynamic spectrum access method based on deep reinforcement learning
CN115190489A