Dynamic spectrum access method, device, equipment and medium
By determining the cognitive vehicle network model and dynamic spectrum access time slot structure in the Internet of Vehicles environment, an optimization model of spectrum access strategy was constructed and the solution was solved using a long and long memory double-duel deep Q network, which solved the problems of low spectrum resource utilization, high energy consumption and low access success rate in the Internet of Vehicles, and an efficient and low energy consumption spectrum access strategy was achieved.
Patent Information
- Application Number
- CN202510308379.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-03-17
AI Technical Summary
In the Internet of Vehicles environment, the traditional static spectrum allocation method cannot adapt to the rapidly changing communication needs, resulting in low utilization of spectrum resources, difficult to ensure communication quality, and high energy consumption and low success rate of spectrum access, making it difficult to meet the high dynamic and complex Internet of Vehicles communication needs.
By determining the cognitive Internet of Vehicles Network Model and dynamic spectrum access time slot structure in the two-lane scenario, a spectrum access strategy optimization model is constructed, and the model is solved using a long-term and short-term memory dual-duel deep Q network to obtain the optimal spectrum access strategy for the secondary user communication link.
While ensuring the efficiency of spectrum access, effectively reduce energy consumption, improve the success rate of spectrum access, realize efficient communication in the Internet of Vehicles, and adapt to dynamic changes in spectrum resources in complex environments.
Smart Images

Figure CN120018146A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of vehicle networking communications, and in particular to a dynamic spectrum access method, device, equipment and medium. Background Art
[0002] With the rapid development of intelligent transportation systems, vehicular ad-hoc networks (VANETs) have attracted widespread attention as the key to realizing vehicle-to-vehicle communications and autonomous driving technologies. Through wireless communications between vehicles, VANETs enable information sharing and collaboration, improving traffic efficiency and safety. However, VANETs need to operate in a complex, changeable, and highly dynamic wireless environment, which makes the scarcity and utilization efficiency of spectrum resources an urgent problem to be solved.
[0003] The traditional static spectrum allocation method cannot adapt to the rapidly changing communication needs of the Internet of Vehicles, resulting in low utilization of spectrum resources and difficulty in ensuring communication quality. Cognitive Radio (CR) technology provides a solution for the Internet of Vehicles to efficiently utilize idle spectrum through dynamic spectrum access (DSA). Cognitive radio can intelligently perceive environmental changes, dynamically adjust spectrum usage strategies, and realize flexible allocation of spectrum resources, thereby improving communication efficiency. Specifically, cognitive radio can dynamically monitor spectrum usage, identify idle spectrum, and allocate these spectrum resources to secondary users (SUs) without interfering with primary users (PUs) through three functional modules: spectrum sensing, spectrum decision-making, and spectrum access. However, how to effectively implement dynamic spectrum access in a highly dynamic Internet of Vehicles environment still faces many technical challenges.
[0004] First, frequent spectrum detection and switching operations will lead to a significant increase in device energy consumption, which not only increases the energy consumption of the vehicle, but also places higher demands on its endurance. In the Internet of Vehicles, vehicles need to continuously monitor the surrounding spectrum environment and respond quickly. This frequent operation will significantly increase the energy consumption of the device. Therefore, how to effectively reduce energy consumption while ensuring spectrum access efficiency is an urgent problem to be solved.
[0005] Secondly, due to the dynamic changes and fierce competition of spectrum resources, vehicles often face access failure when trying to access the spectrum. This situation not only affects the communication quality of the Internet of Vehicles, but may also cause transmission delays or loss of key information. For example, when a traffic accident or emergency occurs, vehicles need to communicate quickly to coordinate actions, but spectrum access failure may result in information not being delivered in time, thus affecting traffic safety and efficiency. Therefore, improving the success rate of spectrum access is another important issue that needs to be solved.
[0006] In addition, the high dynamics and complexity of cognitive vehicle networks result in low channel utilization, making it difficult to ensure high-throughput communication requirements. The high-speed movement of vehicles and frequent changes in network topology make it difficult for traditional static spectrum access methods to adapt to the complex environment in practical applications. When on highways or urban roads, the speed and direction of vehicles change frequently, which places higher requirements on the flexibility and adaptability of spectrum access strategies. Summary of the invention
[0007] The purpose of this application is to provide a dynamic spectrum access method, device, equipment and medium, which can effectively reduce energy consumption, improve the success rate of spectrum access, realize efficient communication of cognitive vehicle networks, and have good adaptability while ensuring spectrum access efficiency.
[0008] To achieve the above objectives, this application provides the following solutions:
[0009] In a first aspect, the present application provides a dynamic spectrum access method, including:
[0010] Determine a cognitive vehicle network model in a two-lane scenario; the cognitive vehicle network model includes: a roadside information unit and multiple vehicle users; the communication between the vehicle user and the roadside information unit is a primary user communication; the communication between the vehicle users is a secondary user communication;
[0011] Determine the dynamic spectrum access time slot structure of the cognitive vehicle network model; each time slot in the dynamic spectrum access time slot structure includes a receiving phase and an access phase; the receiving phase is a phase in which the secondary user communication link receives the spectrum information sent by the roadside information unit; the access phase is a phase in which the secondary user communication link selects a spectrum for access based on the received spectrum information;
[0012] A spectrum access strategy optimization model is constructed based on the dynamic spectrum access time slot structure; the spectrum access strategy optimization model includes: an objective function and a constraint condition; the objective function is constructed with the goal of minimizing the energy consumption and throughput of the secondary user communication link; the constraint condition is determined according to the interference threshold, the energy consumption of the secondary user communication link and the spectrum access time of the secondary user communication link;
[0013] The spectrum access strategy optimization model is solved by using a continuous long short-term memory dual duel deep Q network to obtain the optimal spectrum access strategy for the secondary user communication link; the continuous long short-term memory dual duel deep Q network includes: a long short-term memory network and a dual duel deep Q network connected in sequence; the dual duel deep Q network introduces continuous learning and is constructed based on a dual deep Q network and a duel deep Q network.
[0014] In a second aspect, the present application provides a dynamic spectrum access device, including:
[0015] A network model determination module is used to determine a cognitive vehicle network model in a two-lane scenario; the cognitive vehicle network model includes: a roadside information unit and multiple vehicle users; the communication between the vehicle user and the roadside information unit is a primary user communication; the communication between the vehicle users is a secondary user communication;
[0016] A time slot structure determination module is used to determine the dynamic spectrum access time slot structure of the cognitive vehicle network model; each time slot in the dynamic spectrum access time slot structure includes a receiving phase and an access phase; the receiving phase is a phase in which the secondary user communication link receives the spectrum information sent by the roadside information unit; the access phase is a phase in which the secondary user communication link selects a spectrum for access based on the received spectrum information;
[0017] An access strategy optimization model construction module is used to construct a spectrum access strategy optimization model based on the dynamic spectrum access time slot structure; the spectrum access strategy optimization model includes: an objective function and constraints; the objective function is constructed with the energy consumption and throughput of the secondary user communication link as the goal; the constraints are determined according to the interference threshold, the energy consumption of the secondary user communication link and the spectrum access time of the secondary user communication link;
[0018] The optimization solution module is used to solve the spectrum access strategy optimization model using a continuous long short-term memory dual duel deep Q network to obtain the optimal spectrum access strategy for the secondary user communication link; the continuous long short-term memory dual duel deep Q network includes: a long short-term memory network and a dual duel deep Q network connected in sequence; the dual duel deep Q network introduces continuous learning and is constructed based on a dual deep Q network and a duel deep Q network.
[0019] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the above-mentioned dynamic spectrum access method.
[0020] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon, which implements the above-mentioned dynamic spectrum access method when executed by a processor.
[0021] According to the specific embodiments provided in this application, this application has the following technical effects:
[0022] The present application provides a dynamic spectrum access method, apparatus, device and medium. Each time slot in the dynamic spectrum access time slot structure of the cognitive vehicle network model includes a receiving phase and an access phase. In the time slot structure, the secondary user communication link receives the spectrum information sent by the roadside information unit to select the spectrum for access. The characteristics of the roadside information unit being fixed and not consuming vehicle energy are utilized to improve the problem that different vehicles have different spectrum perception results, effectively alleviate the energy consumption of spectrum perception, and can effectively reduce energy consumption while ensuring spectrum access efficiency. With the goal of minimizing the energy consumption and throughput of the secondary user communication link, a spectrum access strategy optimization model constructed with constraints determined according to the interference threshold, the energy consumption of the secondary user communication link and the spectrum access time as constraints can maximize the throughput under the premise of satisfying the energy consumption and interference constraints, thereby realizing efficient communication in the cognitive vehicle network. The spectrum access strategy optimization model is solved by using a continuous long short-term memory dual duel deep Q network to obtain the optimal spectrum access strategy for the secondary user communication link, which not only improves the success rate of spectrum access, but also effectively overcomes the problems of slow model update and poor adaptability. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0024] Figure 1 This is an application environment diagram of a dynamic spectrum access method in one embodiment of the present application;
[0025] Figure 2 A schematic diagram of a flow chart of a dynamic spectrum access method provided in one embodiment of the present application;
[0026] Figure 3 A schematic diagram of a cognitive Internet of Vehicles network model provided in an embodiment of the present application;
[0027] Figure 4 A schematic diagram of a dynamic spectrum access time slot structure provided in an embodiment of the present application;
[0028] Figure 5A schematic diagram of the model structure of a D3QN network provided in an embodiment of the present application;
[0029] Figure 6 A schematic diagram of the structure of an LSTM network provided in one embodiment of the present application;
[0030] Figure 7 A schematic diagram of a DSA algorithm based on CLD3QN provided in an embodiment of the present application;
[0031] Figure 8 A schematic diagram of functional modules of a dynamic spectrum access device provided by another embodiment of the present application;
[0032] Fig. 9 A schematic diagram of the structure of a computer device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0033] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0034] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.
[0035] The dynamic spectrum access method provided in the embodiment of the present application can be applied to Figure 1In the application environment shown. Among them, the terminal 102 communicates with the server 104 through the network. The data storage system can store data that the server 104 needs to process. The data storage system can be set up separately, integrated on the server 104, or placed on the cloud or other servers. The terminal 102 can send the cognitive Internet of Vehicles network model to the server 104. After the server 104 receives the cognitive Internet of Vehicles network model, for the cognitive Internet of Vehicles network model, the server 104 determines the dynamic spectrum access time slot structure of the cognitive Internet of Vehicles network model; constructs a spectrum access strategy optimization model based on the dynamic spectrum access time slot structure; and uses a continuous long and short-term memory dual duel deep Q network to solve the spectrum access strategy optimization model to obtain the optimal spectrum access strategy for the secondary user communication link. The server 104 can feedback the obtained optimal spectrum access strategy for the secondary user communication link to the terminal 102. In addition, in some embodiments, the dynamic spectrum access method can also be implemented independently by the server 104 or the terminal 102. For example, the terminal 102 can directly process the cognitive Internet of Vehicles network model, or the server 104 can obtain the cognitive Internet of Vehicles network model from the data storage system and process the cognitive Internet of Vehicles network model.
[0036] The terminal 102 may be, but is not limited to, various desktop computers, laptop computers, smart phones, tablet computers, IoT devices, and portable wearable devices. The IoT devices may be smart speakers, smart TVs, smart air conditioners, smart vehicle-mounted devices, etc. The portable wearable devices may be smart watches, smart bracelets, head-mounted devices, etc. The server 104 may be implemented as an independent server or a server cluster consisting of multiple servers, or may be a cloud server.
[0037] In an exemplary embodiment, Figure 2 As shown, a dynamic spectrum access method is provided, which is executed by a computer device, and can be executed by a computer device such as a terminal or a server alone, or by a terminal and a server together. In the embodiment of the present application, the method is applied to Figure 1 The server 104 in the example is used as an example to illustrate the method, which includes the following steps 201 to 204. Among them:
[0038] Step 201, determine a cognitive Internet of Vehicles network model in a two-lane scenario.
[0039] Among them, the cognitive Internet of Vehicles network model includes: a road side information unit (Road Side Unit, RSU) and multiple vehicle users; the communication between the vehicle user and the road side information unit is primary user communication; and the communication between the vehicle users is secondary user communication.
[0040] Step 202: determine the dynamic spectrum access time slot structure of the cognitive Internet of Vehicles network model.
[0041] Among them, each time slot in the dynamic spectrum access time slot structure includes a receiving stage and an access stage; the receiving stage is the stage in which the secondary user communication link receives the spectrum information sent by the roadside information unit; the access stage is the stage in which the secondary user communication link selects the spectrum for access based on the received spectrum information.
[0042] Step 203: construct a spectrum access strategy optimization model based on the dynamic spectrum access time slot structure.
[0043] Among them, the spectrum access strategy optimization model includes: an objective function and constraints; the objective function is constructed with the goal of minimizing the energy consumption and throughput of the secondary user communication link; the constraints are determined based on the interference threshold, the energy consumption of the secondary user communication link and the spectrum access time of the secondary user communication link.
[0044] Step 204: Use a continuous long short-term memory double dueling deep Q-network (Continual LSTM DoubleDueling Deep Q-Network, CLD3QN) to solve the spectrum access strategy optimization model to obtain the optimal spectrum access strategy for the secondary user communication link.
[0045] Among them, the continuous long short-term memory dual dueling deep Q network includes: a long short-term memory network and a dual dueling deep Q network connected in sequence; the dual dueling deep Q network introduces continuous learning (Continual Learninginng, CL) and is constructed based on a double deep Q network (Double Deep Q-Network, DDQN) and a dueling deep Q network (Dueling DQN).
[0046] The implementation of the above steps 201 to 204 can effectively reduce energy consumption, improve the success rate of spectrum access, and achieve efficient communication of cognitive vehicle networks while ensuring spectrum access efficiency, and has good adaptability.
[0047] In another exemplary embodiment of the present application, the cognitive Internet of Vehicles network model in the two-lane scenario in step 201 and the dynamic spectrum access time slot structure of the cognitive Internet of Vehicles network model are further introduced.
[0048] Considering the complexity of future networks, a cognitive IOV network with multiple vehicle users and roadside information units is designed in an urban environment. The cognitive IOV network model is as follows: Figure 3As shown, this embodiment realizes the vehicle network communication in a two-lane scenario in an urban environment, including a base station (BS), a roadside information unit RSU and a vehicle user. The communication modes mainly include communication between the roadside information unit and the vehicle user (V2I) and communication between the vehicle user and the vehicle user (V2V). In the model, V2V is defined as secondary user communication and V2I is defined as primary user communication. In order to facilitate effective communication and collision-free transmission, each primary user communication is assigned a separate radio channel. Figure 3 T stands for time slot, R stands for receive, and A stands for access.
[0049] Assume that the set of authorized frequency bands for V2I communication is l∈{1,2,…,L}, and the set of frequency bands for V2V communication is n∈{1,2,…,N}, where L and N represent the number of V2I and V2V communications, respectively. In addition, assuming that the number of spectrums in the current Internet of Vehicles environment is the same as the number of authorized spectrums, the channel combination in the current system model is m∈{1,2,…,M}, and the cross-channel interference can be ignored. In order to improve the utilization of spectrum resources, the secondary user communication opportunity is allowed to access the idle frequency bands authorized by the primary user but not in use.
[0050] like Figure 4 As shown in part (a), in the traditional dynamic spectrum access (DSA) time slot structure, the V2V link needs to sense spectrum holes by itself and then choose to access the idle spectrum. Such a time slot structure is difficult to meet the requirements of different vehicle communication perceptions, and will waste too much energy consumption of the V2V link for frequent spectrum perception. This embodiment uses the characteristics of the roadside information unit being fixed and not consuming vehicle energy to improve the problem of different vehicles having different spectrum perception results, and effectively alleviates the energy consumption of spectrum perception.
[0051] Therefore, this embodiment proposes Figure 4 In the improved DSA time slot structure shown in part (b), the V2V link will receive the spectrum information sent by the roadside information unit before spectrum access, and then select the appropriate spectrum for access. In this process, the access time τ of each V2V link is limited to the current time slot, that is, it satisfies the constraint condition of τ≥0. In addition, in order to ensure the rational use of spectrum resources, the access time τ should be less than the total time slot T, that is, it satisfies the constraint condition T-τ≥0, where T represents the duration of the entire time slot. This design ensures that each V2V link can complete the access operation within the limited time and avoids communication conflicts or resource waste due to exceeding the time slot time.
[0052] In another exemplary embodiment of the present application, the idea of constructing a spectrum access strategy optimization model in step 203 is introduced.
[0053] According to the improved dynamic spectrum access time slot structure, the calculation formula of the energy consumption of the secondary user communication link n on the channel m can be expressed as:
[0054]
[0055] Where τ represents the spectrum access time of the secondary user communication link, that is, the time consumed by access, P n represents the transmission power of the secondary user communication link n, represents the channel gain of the secondary user communication link n on channel m. At the same time, in order to optimize energy consumption, the secondary user communication link must also meet the maximum energy consumption constraint during the access process to ensure that the system controls energy consumption while maximizing throughput, expressed as
[0056] In the considered vehicle networking environment, the propagation condition between communication users is assumed to be line of sight (LOS), and the channel model is assumed to be a Rician fading channel. Therefore, the channel gain of the secondary user communication link n on channel m is The calculation formula is as follows:
[0057]
[0058] Where k represents the ratio of the received signal power to the power variance of the scattering path under LOS, and θ represents the phase of the received signal. CN(·) is a circularly symmetric complex Gaussian random variable, is the scale parameter of the secondary user communication link n. The path loss of the secondary user communication link n Communication distance and the carrier frequency f c Related, the calculation formula of the path loss of the secondary user communication link n is as follows:
[0059]
[0060] in, represents the communication distance of the secondary user communication link n, f c Indicates the carrier frequency of the wireless channel, PL α , PL β and PL γ represent the path loss of the reference distance, the path loss index and the frequency correlation of the path loss respectively, and GHz represents the unit of the carrier frequency of the line channel.
[0061] In the model of this embodiment, it is assumed that the RSU and the vehicle's transceiver both use one antenna. Represent the location information of RSUi and vehicle j, where i∈{1,2,…,I}, j∈{1,2,…,J}. Therefore, the communication distance between another vehicle j' and vehicle j in the secondary user communication link n is calculated. The formula is as follows:
[0062]
[0063] Indicates the position information of vehicle j'.
[0064] At the same time, the primary user communication link l is used as the primary user communication, and its channel gain on channel m is The calculation method is as follows:
[0065]
[0066] in, is the scale parameter of the primary user communication link l. The path loss of the primary user communication link l Communication distance and the carrier frequency f c Related, the formula is as follows:
[0067]
[0068] in, The communication distance between the primary user communication link l and the vehicle j is The calculation formula is as follows:
[0069]
[0070] After considering the improved dynamic spectrum access time slot structure and its energy consumption constraints, this embodiment focuses on the impact of interference on the access of secondary user communication links. It can be expressed as:
[0071]
[0072] in, represents the strategy of the current secondary user communication link access channel m, B m represents the bandwidth of channel m, Represents the transmission power of the secondary user communication link n, and will be based on the access strategy Change the size, represents the noise power spectrum density. According to formula (8), the access strategy The impact on throughput is significant.
[0073] Therefore, this embodiment designs a new hybrid access mode that combines the characteristics of the Overlay and Underlay access modes. After receiving the channel information sent by the RSU, the SU makes a decision to access the channel m based on the channel status information. If there is any information that is perceived as an error, the SU determines whether the current channel exceeds the interference threshold I of the current channel. th , when it is lower than , the transmission power is reduced to complete the access process, otherwise the transmission fails. The situation where the secondary user communication link accesses channel m under different interferences, that is, the spectrum access strategy of the secondary user communication link accessing channel m The expression is:
[0074]
[0075] in, It represents the interference of other primary user communication link l on secondary user communication link n on channel m at the current time t. The calculation formula is as follows:
[0076]
[0077] Respectively represent the interference threshold of a primary user communication link, the interference threshold of a secondary user communication link, and the maximum interference threshold. When there is no interference on the current channel m, the channel can be successfully accessed. When there is a primary user communication link on channel m and the interference is less than the interference threshold on the primary user communication link When the secondary user communication link is occupied on channel m and the interference is less than the interference threshold on the secondary user communication link, When , the secondary user communication link can successfully access the current channel without reducing power; otherwise, it needs to reduce power to ensure successful access to the channel. If there is interference, access fails.
[0078] Therefore, the optimization problem can be expressed as the secondary user communication link n looking for the best access strategy among all authorized channels m, while ensuring the maximum throughput of the secondary user communication link and the lowest energy consumption, and jointly optimizing the access time and interference threshold to construct a spectrum access strategy optimization model.
[0079] Based on the above content, step 203 specifically includes:
[0080] (1) Determine the number of spectrum based on the dynamic spectrum access time slot structure; each spectrum segment corresponds to a channel.
[0081] (2) Calculate the communication distance between two vehicle users in the secondary user communication link. The calculation formula is shown in formula (4).
[0082] (3) The path loss of the secondary user communication link is calculated according to the communication distance between the two vehicle users in the secondary user communication link and the carrier frequency. The calculation formula is shown in formula (3).
[0083] (4) The channel gain of the secondary user communication link on each channel is calculated according to the path loss of the secondary user communication link. The calculation formula is shown in formula (2).
[0084] (5) According to the channel gain of the secondary user communication link on each channel, the transmission power of the secondary user communication link and the spectrum access time of the secondary user communication link, the energy consumption of the secondary user communication link on each channel is calculated. The calculation formula is shown in formula (1).
[0085] (6) Calculate the communication distance between the vehicle user and the roadside information unit in the primary user communication link. The calculation formula is shown in formula (7).
[0086] (7) The path loss of the primary user communication link is calculated based on the communication distance between the vehicle user and the roadside information unit in the primary user communication link and the carrier frequency. The calculation formula is shown in formula (6).
[0087] (8) The channel gain of the primary user communication link on each channel is calculated according to the path loss of the primary user communication link. The calculation formula is shown in formula (5).
[0088] (9) The interference of the primary user communication link on the secondary user communication link on each channel is calculated according to the channel gain of the primary user communication link on each channel. The calculation formula is shown in formula (10).
[0089] (10) The spectrum access strategy for the secondary user communication link to access each channel is determined according to the interference of the primary user communication link to the secondary user communication link and the interference threshold. The calculation formula is shown in formula (9).
[0090] (11) According to the spectrum access strategy of the secondary user communication link accessing each channel, the bandwidth of the channel, the spectrum access time of the secondary user communication link, the channel gain of the primary user communication link on each channel, the transmission power and noise power spectrum density of the secondary user communication link, the throughput of the secondary user communication link when accessing each channel is calculated. The calculation formula of the throughput is shown in formula (8).
[0091] (12) A spectrum access strategy optimization model is constructed based on the energy consumption of the secondary user communication link on each channel, the throughput of the secondary user communication link when accessing each channel, the interference threshold and the spectrum access time of the secondary user communication link.
[0092] The spectrum access strategy optimization model is expressed as:
[0093]
[0094] Among them, max means taking the maximum value; st means the constraint condition; N means the number of secondary user communication links; M means the number of channels; represents the throughput of secondary user communication link n when accessing channel m; I represents the spectrum access strategy of the secondary user communication link access channel m; th represents the interference threshold; τ represents the spectrum access time of the secondary user communication link; represents the energy consumption of secondary user communication link n on channel m; represents the interference threshold of the secondary user communication link; represents the interference threshold of the primary user communication link; represents the maximum interference threshold; T represents the total time to complete the access process. In this embodiment, Characterize the different strategies of the current secondary user communication link access channel m, and the energy consumption of the secondary user communication link in the designed hybrid access mode There is a maximum value constraint, interference threshold There are minimum and maximum value constraints.
[0095] Formula (11) is a nonlinear optimization problem with multiple constraints, and it is difficult to solve it directly. Due to the dynamic changes in the Internet of Vehicles environment, vehicles face different interference thresholds and energy consumption constraints when making DSA decisions, and it is difficult to obtain the optimal spectrum access strategy in real time. In addition, in high-density traffic scenarios, competition between vehicles for limited spectrum resources may cause serious interference and reduce the overall transmission efficiency. Therefore, the focus is on designing an enhanced dynamic spectrum access algorithm (step 204) to maximize throughput while meeting energy consumption and interference constraints, and achieve efficient communication in cognitive Internet of Vehicles.
[0096] In another exemplary embodiment of the present application, the dynamic spectrum access algorithm based on CLD3QN in step 204 is introduced. For the convenience of introduction, subsequent secondary user communication is represented by V2V, and primary user communication is represented by V2I.
[0097] exist Figure 3 In the system model of , multiple V2V links attempt to access the limited spectrum occupied by the V2I link, which can be modeled as a multi-agent RL problem. Each V2V link acts as an agent, interacting with the unknown communication environment to gain experience, which is then used to guide its own strategy design. Multiple V2V agents jointly explore the environment and improve the spectrum access strategy based on their own observations of the environment state.
[0098] In this embodiment, CLD3QN is used to implement distributed DSA of V2V links in dynamic vehicle environments. LSTM is mainly used to save the historical spectrum decision data of V2V, D3QN is mainly used to determine the optimal spectrum access strategy of V2V, and CL is to adapt to new changes. The V2V link first obtains the channel state from the RSU, selects the available channel for data transmission, and obtains the corresponding reward. Then, the channel state, the action at the current moment, and the reward are input into the LSTM network. Finally, the calculation results of the LSTM network are input into the network of D3QN to estimate the optimal channel selection strategy.
[0099] 1) D3QN network.
[0100] D3QN is an improvement on the Dual Deep Q Network (DDQN) and Dueling Deep Q Network (Dueling DQN). D3QN decomposes the Q value function into two parts: state value and advantage, making the valuation more accurate and improving the learning effect. The model structure of the D3QN network is as follows: Figure 5 Dynamic spectrum access in the cognitive Internet of Vehicles environment involves the time-varying and spatial variability of spectrum resources. This method is particularly suitable for complex environments of dynamic spectrum access and can effectively improve spectrum utilization.
[0101] D3QN uses two deep neural networks for data learning, namely the Evaluation Q-Network (EQN) and the Target Q-Network (TQN). The Evaluation Q-Network obtains the current state and action samples from the experience pool and predicts the Q value of the specific action. The Target Q-Network obtains the next state sample from the experience pool and predicts the best Q value of all actions starting from that state. The Evaluation Q-Network and the Target Q-Network are the same in network structure, and when solving the Q-value function, they are divided into the state value function and the advantage function to better estimate the Q value.
[0102] In D3QN, both the evaluation Q network and the target Q network adopt the duel network structure, decomposing the Q value function into the state value function V(s,ω) and the advantage function A(s,a,ω). Among them, s is the current state, a is the action, and ω is the network weight. The Q value function can be expressed as:
[0103]
[0104] Where |A| is the size of the action set, ω is the parameter of the evaluation Q network, and α is the parameter of the equilibrium state value function and the action advantage function.
[0105] The target Q value update formula in D3QN is:
[0106]
[0107] Where ω' is the parameter of the target Q network.
[0108] The Q-value update strategy in D3QN can be described as:
[0109]
[0110] Where r is the reward of the action chosen for the current state, γ∈(0,1) is the Q-value discount rate of D3QN, and β∈(0,1) is the learning rate.
[0111] The selected action is determined by the ∈-greedy strategy as follows:
[0112]
[0113] The D3QN algorithm uses an experience replay mechanism to store the experience transfer samples obtained by the agent from the environment at each moment. After executing several steps, a batch size of samples is randomly drawn from the experience pool. Then, D3QN uses a small batch stochastic semi-gradient descent method to update the network parameters ω to minimize the loss function L(ω), which is expressed as follows:
[0114]
[0115] During the training process, the weights of EQN are updated through gradient descent and back-propagation techniques. The target network TQN remains unchanged. After several training rounds, the weights of EQN are copied to TQN. This helps TQN to obtain improved weights and thus also predict more accurate target Q values.
[0116] 2) LSTM network.
[0117] In order to learn efficient spectrum access strategies based on historical experience, this embodiment uses an LSTM network to save long-term historical spectrum selection results and solve the problems of gradient vanishing and gradient explosion during long sequence training. In the process of reinforcement learning, a large amount of historical information can help vehicles quickly understand the characteristics of cognitive vehicle networking environments. Figure 6 As shown in the figure, the LSTM network consists of three gates. The first is the Forget Gate, which determines which information in the memory cell needs to be forgotten. Its calculation formula is:
[0118] f t =σ(W f ·[h t-1 ,x t ]+b f )(twenty two)
[0119] Among them, f t is the output of the forget gate and f t ∈(0, 1), σ is the Sigmoid function, W f is the weight matrix of the forget gate, h t-1 is the hidden state at the previous moment, x t is the input at the current moment, b f is the bias term of the forget gate. The second is the input gate, which determines how much new information is added to the memory cell. Its calculation formula is:
[0120] i t =σ(W i ·[h t-1 ,x t ]+b i )(twenty three)
[0121] Among them, i t is the output of the input gate, W i is the weight matrix of the input gate, b i is the bias term of the input gate; the candidate value of the new information is generated through a tanh layer:
[0122] C' t =tanh(W C ·[h t-1 ,x t ]+b C )(twenty four)
[0123] Among them, C' t is the new candidate memory, tanh is the hyperbolic tangent function, W C represents the weight matrix of candidate memory, b C It is the bias item of the candidate memory; after obtaining the candidate value, the memory cell (Cell State) will be updated. The state of the memory cell is updated by combining the forget gate and the input gate:
[0124] C t =f t ·C t-1 +i t ·C' t (25)
[0125] Among them, C t is the current state of the memory cell, C t-1 The state of the memory cell at the previous moment. The third is the output gate, which determines the hidden state at the current moment. The calculation formula is:
[0126] o t=σ(W o ·[h t-1 ,x t ]+b o )(26)
[0127] Among them, t is the output of the output gate, W o is the weight matrix of the output gate, b o is the bias term of the output gate; the hidden state is obtained by combining the memory cell state and the output gate:
[0128] h t =o t tanh(C t )(27)
[0129] Among them, h t Indicates the hidden state at the current moment.
[0130] In summary, input x t and the hidden state h at the previous moment t-1 and memory cell state C t-1 After calculation of the above formula, the hidden state h at the current moment is output t and memory cell state C t .
[0131] 3) Dynamic spectrum access (DSA) algorithm based on CLD3QN.
[0132] In order to realize intelligent spectrum access for cognitive Internet of Vehicles, this embodiment proposes a dynamic spectrum access algorithm for cognitive Internet of Vehicles based on CLD3QN. LSTM networks are good at processing time series data and can capture historical information and time dependence of spectrum usage, while CL can adapt to dynamic environments more efficiently. By using two Q networks (evaluation Q network and target Q network) to separate action selection and Q value update, the problem of Q value over-estimation can be effectively reduced and the stability of learning can be improved. The Q value function is decomposed into state value function and advantage function in the two Q networks, which can more accurately estimate the value of the state and the relative advantages of different actions. This decomposition method can reduce the variance of Q value estimation and improve the stability and accuracy of decision-making. By combining CL and LSTM, the model can adapt to changes in dynamic spectrum environments and continuously update and improve spectrum access strategies. After adding the idea of continuous learning, the new loss function changes to:
[0133]
[0134] Among them, L CLD3QN (ω') represents the loss function; represents the expected value of data sampled from the experience replay pool; Q e(s, a; ω) represents the Q value of the Q network output; Q t (s,a) represents the Q value of the target Q network output; ψ∥ω'-ω∥ 2 represents L2 regularization (also known as weight decay), which is used to prevent model overfitting; ψ represents the regularization parameter; s represents the channel state; a represents the action; r represents the reward; s' represents the next channel state; ω represents the evaluation Q network parameters of the previous batch, which is used to constrain the changes of new and old parameters and avoid catastrophic forgetting; ω' represents the current dual duel deep Q network parameters.
[0135] based on Figure 4 The time slot structure shown in the figure proposes a coverage spectrum access scheme based on CLD3QN. In DSA, the V2V link can only use the idle channels not occupied by the V2I link. If a channel is being used by V2I, V2I cannot access the channel in any way. In time slot T, the vehicle obtains the status of M channels through the broadcast of RSU. Using binary variables Represents the state of channel m, then the channel state obtained by the V2V link can be expressed by the vector Indicates that Indicates the busy state of the channel in the current period. Indicates that the channel is in an idle state during the current period, and Represents the interference threshold of the current channel. Assume that the vehicle selects only one channel for communication in each time slot. After obtaining the channel status, the V2V link takes the action of accessing a channel. in Indicates that the V2V link selects channel m for access in the current period. It indicates that the V2V link does not select access to channel m in the current period.
[0136] When a V2V link successfully enters a channel and completes transmission within a time slot, it will receive a q-learning reward. t and action A t The reward R t It is expressed as:
[0137]
[0138] Among them, Γ m , They represent the throughput, energy consumption and interference of V2V link n in the current channel m respectively, χ is the energy consumption coefficient, and ι is the interference coefficient.
[0139] Therefore, under different access conditions, the reward function of the V2V link can be expressed as:
[0140]
[0141] Among them, R t (S t ,A t ) represents the reward at the current time t for the channel state and the action at the current time t; represents the throughput of secondary user communication link n when accessing channel m; represents the energy consumption of the secondary user communication link n on the channel m; χ represents the energy consumption coefficient; ι represents the interference coefficient; represents the interference of the primary user communication link on the secondary user communication link n on the channel m; represents the interference threshold of the primary user communication link; represents the interference threshold of the secondary user communication link; C represents a constant; It represents the interference of the primary user communication link on the secondary user communication link n on the channel m at the current time t.
[0142] The DSA algorithm based on CLD3QN proposed in this embodiment is as follows Figure 7 As shown. The LSTM network stores the long-term historical observation data perceived by the RSU, and the D3QN calculates the optimal spectrum access strategy for the V2V link, updating the continuous learning every 10 steps to adapt to the dynamic spectrum environment. For ease of understanding, this embodiment illustrates a simple DSA process in the algorithm process. Figure 7 As shown in the figure, RSU senses the channel status from the environment and broadcasts the channel status to all vehicle users. V2VAgent1 selects to access channel 2 based on the received information. The perceived status of channel 2 is "0", indicating that channel 2 is available; the actual status is "1", indicating that channel 3 is occupied by V2I.
[0143] Therefore, V2VAgent2 collides with V2I after accessing channel 3. If the interference does not exceed Rewards Otherwise the reward is V2VAgent2 selects to access channel 3 based on the received information. The perceived state of channel 3 is "0", indicating that channel 3 is available; the actual state is "0", indicating that channel 3 is not occupied by V2I. Next, V2VAgent2 successfully accesses channel 3 and the reward is At the same time, V2VAgentN also chooses to access channel 3 according to the received information. As mentioned above, V2VAgentN also chooses channel 3 to access. Therefore, after V2VAgent3 accesses channel 2, it collides with V2VAgent1. The reward is Otherwise the reward is When more than two users in a channel, that is, a third user attempts to access the channel, it is defined as causing serious interference to all three users and destroying the communication quality. The reward at this time is -C, where C is a constant. The reward is 0 when the agent does not select a channel.
[0144] The pseudo code of the DSA algorithm based on CLD3QN is described as follows:
[0145] Algorithm 1: DSA algorithm based on CLD3QN
[0146] Input: Initialize the CLD3QN network of each SU, set the learning rate β, discount factor γ, action selection strategy and false alarm probability Pf, and initialize the parameters α of the continuous learning model.
[0147] Output: The optimal spectrum access strategy for each SU.
[0148]
[0149] Based on the above introduction, step 204 specifically includes:
[0150] (1) Obtain the channel status perceived by the roadside information unit from the environment.
[0151] (2) Based on the spectrum access strategy optimization model, when the secondary user communication link selects a channel according to the received channel state and completes data transmission within a time slot, a reward is calculated under the current channel state and the current action. The calculation formula of the reward under the current channel state and the current action is shown in formula (30).
[0152] (3) The channel state at the current moment, the action at the current moment, and the reward are input into the long short-term memory network to obtain the channel state at the next moment.
[0153] (4) Constructing a dual duel deep Q network; the dual duel deep Q network includes: an evaluation Q network and a target Q network.
[0154] (5) Inputting the channel state and action at the current moment into the evaluation Q network, inputting the channel state at the next moment into the target Q network, training the dual duel deep Q network with the goal of minimizing the loss function, and using the spectrum access strategy corresponding to the Q value output by the trained target Q network as the optimal spectrum access strategy for the secondary user communication link; the loss function is determined by introducing continuous learning based on the reward, the Q value output by the evaluation Q network, and the Q value output by the target Q network. The expression of the loss function is shown in formula (28).
[0155] The dynamic spectrum access method of this application is a green and efficient dynamic spectrum access method based on the continuous long short-term memory dual dueling deep Q network (CLD3QN), which successfully combines CR with VANETs and significantly improves the utilization rate of spectrum resources. By introducing a hybrid access mode and comprehensively utilizing Overlay and Underlay technologies, the access success rate performance of the system in scenarios with different numbers of users is optimized. The proposed CLD3QN algorithm combines continuous learning and long short-term memory networks, effectively overcoming the problems of slow model updates and poor adaptability in traditional methods.
[0156] Based on the same inventive concept, the embodiment of the present application also provides a dynamic spectrum access device for implementing the dynamic spectrum access method involved above. The implementation solution provided by the device to solve the problem is similar to the implementation solution recorded in the above method, so the specific limitations in one or more dynamic spectrum access device embodiments provided below can refer to the limitations of the dynamic spectrum access method above, and will not be repeated here.
[0157] In an exemplary embodiment, Figure 8 As shown, a dynamic spectrum access device is provided, including:
[0158] The network model determination module 801 is used to determine the cognitive vehicle network model in a two-lane scenario; the cognitive vehicle network model includes: a roadside information unit and multiple vehicle users; the communication between the vehicle user and the roadside information unit is primary user communication; the communication between the vehicle users is secondary user communication.
[0159] The time slot structure determination module 802 is used to determine the dynamic spectrum access time slot structure of the cognitive vehicle network model; each time slot in the dynamic spectrum access time slot structure includes a receiving phase and an access phase; the receiving phase is a phase in which the secondary user communication link receives the spectrum information sent by the roadside information unit; the access phase is a phase in which the secondary user communication link selects the spectrum for access based on the received spectrum information.
[0160] The access strategy optimization model construction module 803 is used to construct a spectrum access strategy optimization model based on the dynamic spectrum access time slot structure; the spectrum access strategy optimization model includes: an objective function and constraints; the objective function is constructed with the goal of minimizing the energy consumption and throughput of the secondary user communication link; the constraints are determined based on the interference threshold, the energy consumption of the secondary user communication link and the spectrum access time of the secondary user communication link.
[0161] The optimization solution module 804 is used to solve the spectrum access strategy optimization model using a continuous long short-term memory dual duel deep Q network to obtain the optimal spectrum access strategy for the secondary user communication link; the continuous long short-term memory dual duel deep Q network includes: a long short-term memory network and a dual duel deep Q network connected in sequence; the dual duel deep Q network introduces continuous learning and is constructed based on a dual deep Q network and a duel deep Q network.
[0162] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal. The internal structure diagram thereof may be as follows: Fig. 9 As shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, referred to as I / O) and a communication interface. Among them, the processor, the memory and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store a cognitive vehicle network model in a two-lane scenario. The input / output interface of the computer device is used to exchange information between the processor and an external device. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a dynamic spectrum access method is implemented.
[0163] Those skilled in the art will understand that Fig. 9 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components. In an exemplary embodiment, a computer device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.
[0164] In an exemplary embodiment, a computer-readable storage medium is provided, storing a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0165] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0166] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.
[0167] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM may be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM).
[0168] The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. The non-relational database may include a distributed database based on blockchain, etc., but is not limited thereto. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, etc., but is not limited thereto.
[0169] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0170] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.
Claims
1. A dynamic spectrum access method, characterized in that: The dynamic spectrum access method comprises: Determine a cognitive vehicle network model in a two-lane scenario; the cognitive vehicle network model includes: a roadside information unit and multiple vehicle users; the communication between the vehicle user and the roadside information unit is a primary user communication; the communication between the vehicle users is a secondary user communication; Determine the dynamic spectrum access time slot structure of the cognitive vehicle network model; each time slot in the dynamic spectrum access time slot structure includes a receiving phase and an access phase; the receiving phase is a phase in which the secondary user communication link receives the spectrum information sent by the roadside information unit; the access phase is a phase in which the secondary user communication link selects a spectrum for access based on the received spectrum information; A spectrum access strategy optimization model is constructed based on the dynamic spectrum access time slot structure; the spectrum access strategy optimization model includes: an objective function and a constraint condition; the objective function is constructed with the goal of minimizing the energy consumption and throughput of the secondary user communication link; the constraint condition is determined according to the interference threshold, the energy consumption of the secondary user communication link and the spectrum access time of the secondary user communication link; The spectrum access strategy optimization model is solved by using a continuous long short-term memory dual duel deep Q network to obtain the optimal spectrum access strategy for the secondary user communication link; the continuous long short-term memory dual duel deep Q network includes: a long short-term memory network and a dual duel deep Q network connected in sequence; the dual duel deep Q network introduces continuous learning and is constructed based on a dual deep Q network and a duel deep Q network.
2. The dynamic spectrum access method according to claim 1, characterized in that: Building a spectrum access strategy optimization model based on the dynamic spectrum access time slot structure specifically includes: Determining the number of spectrum based on the dynamic spectrum access time slot structure; each spectrum corresponds to a channel; Calculate the communication distance between two vehicle users in the secondary user communication link; Calculate the path loss of the secondary user communication link according to the communication distance between two vehicle users in the secondary user communication link and the carrier frequency; Calculate the channel gain of the secondary user communication link on each channel according to the path loss of the secondary user communication link; Calculate the energy consumption of the secondary user communication link on each channel according to the channel gain of the secondary user communication link on each channel, the transmission power of the secondary user communication link and the spectrum access time of the secondary user communication link; Calculating the communication distance between the vehicle user and the roadside information unit in the primary user communication link; Calculate the path loss of the primary user communication link according to the communication distance between the vehicle user and the roadside information unit in the primary user communication link and the carrier frequency; Calculate the channel gain of the primary user communication link on each channel according to the path loss of the primary user communication link; Calculate the interference of the primary user communication link to the secondary user communication link on each channel according to the channel gain of the primary user communication link on each channel; Determine a spectrum access strategy for the secondary user communication link to access each channel according to the interference of the primary user communication link to the secondary user communication link and the interference threshold; Calculate the throughput of the secondary user communication link when accessing each channel according to the spectrum access strategy of the secondary user communication link accessing each channel, the bandwidth of the channel, the spectrum access time of the secondary user communication link, the channel gain of the primary user communication link on each channel, the transmission power and noise power spectral density of the secondary user communication link; A spectrum access strategy optimization model is constructed based on the energy consumption of the secondary user communication link on each channel, the throughput of the secondary user communication link when accessing each channel, the interference threshold and the spectrum access time of the secondary user communication link.
3. The dynamic spectrum access method according to claim 2, characterized in that: The spectrum access strategy optimization model is expressed as: T-τ≥0 τ>0 Among them, max means taking the maximum value; st means the constraint condition; N means the number of secondary user communication links; M means the number of channels; represents the throughput of secondary user communication link n when accessing channel m; I represents the spectrum access strategy of the secondary user communication link access channel m; th represents the interference threshold; τ represents the spectrum access time of the secondary user communication link; represents the energy consumption of secondary user communication link n on channel m; represents the interference threshold of the secondary user communication link; represents the interference threshold of the primary user communication link; represents the maximum interference threshold; T represents the total time to complete the access process.
4. The dynamic spectrum access method according to claim 1, characterized in that: The spectrum access strategy optimization model is solved by using a persistent long short-term memory dual dueling deep Q network to obtain the optimal spectrum access strategy for the secondary user communication link, which specifically includes: Obtaining the channel state perceived by the roadside information unit from the environment; Based on the spectrum access strategy optimization model, when the secondary user communication link selects a channel according to the received channel state and completes data transmission within a time slot, a reward is calculated under the channel state at the current moment and the action at the current moment; Inputting the current channel state, the current action and the reward into the long short-term memory network to obtain the channel state at the next moment; Constructing a dual duel deep Q network; the dual duel deep Q network includes: an evaluation Q network and a target Q network; The channel state at the current moment and the action at the current moment are input into the evaluation Q network, and the channel state at the next moment is input into the target Q network. The dual duel deep Q network is trained with the goal of minimizing the loss function, and the spectrum access strategy corresponding to the Q value output by the trained target Q network is used as the optimal spectrum access strategy for the secondary user communication link; the loss function is determined by introducing continuous learning based on the reward, the Q value output by the evaluation Q network, and the Q value output by the target Q network.
5. The dynamic spectrum access method according to claim 4, characterized in that: The calculation formula for the reward under the current channel state and the current action is: Among them, R t (S t ,A t ) represents the reward at the current time t for the channel state and the action at the current time t; represents the throughput of secondary user communication link n when accessing channel m; represents the energy consumption of the secondary user communication link n on the channel m; χ represents the energy consumption coefficient; ι represents the interference coefficient; represents the interference of the primary user communication link on the secondary user communication link n on the channel m; represents the interference threshold of the primary user communication link; represents the interference threshold of the secondary user communication link; C represents a constant; It represents the interference of the primary user communication link on the secondary user communication link n on the channel m at the current time t.
6. The dynamic spectrum access method according to claim 4, characterized in that: The expression of the loss function is: Among them, L CLD3QN (ω') represents the loss function; represents the expected value of data sampled from the experience replay pool; Q e (s, a; ω) represents the Q value of the Q network output; Q t (s,a) represents the Q value of the target Q network output; ψ∥ω'-ω∥ 2 represents L2 regularization; ψ represents the regularization parameter; s represents the channel state; a represents the action; r represents the reward; s' represents the next channel state; ω represents the evaluation Q network parameters of the previous batch; ω' represents the current dual duel deep Q network parameters.
7. The dynamic spectrum access method according to claim 3, characterized in that: The calculation formula for the energy consumption of secondary user communication link n on channel m is: Among them, P n represents the transmission power of the secondary user communication link n; represents the channel gain of secondary user communication link n on channel m; The calculation formula for the throughput of secondary user communication link n when accessing channel m is: Among them, B m represents the bandwidth of channel m; represents the transmission power of the secondary user communication link n; represents the noise power spectral density; The spectrum access strategy of the secondary user communication link access channel m is expressed as: in, It represents the interference of the primary user communication link on the secondary user communication link n on the channel m at the current time t.
8. A dynamic spectrum access device, characterized in that: The dynamic spectrum access device comprises: A network model determination module is used to determine a cognitive vehicle network model in a two-lane scenario; the cognitive vehicle network model includes: a roadside information unit and multiple vehicle users; the communication between the vehicle user and the roadside information unit is a primary user communication; the communication between the vehicle users is a secondary user communication; A time slot structure determination module is used to determine the dynamic spectrum access time slot structure of the cognitive vehicle network model; each time slot in the dynamic spectrum access time slot structure includes a receiving phase and an access phase; the receiving phase is a phase in which the secondary user communication link receives the spectrum information sent by the roadside information unit; the access phase is a phase in which the secondary user communication link selects a spectrum for access based on the received spectrum information; An access strategy optimization model construction module is used to construct a spectrum access strategy optimization model based on the dynamic spectrum access time slot structure; the spectrum access strategy optimization model includes: an objective function and constraints; the objective function is constructed with the energy consumption and throughput of the secondary user communication link as the goal; the constraints are determined according to the interference threshold, the energy consumption of the secondary user communication link and the spectrum access time of the secondary user communication link; The optimization solution module is used to solve the spectrum access strategy optimization model using a continuous long short-term memory dual duel deep Q network to obtain the optimal spectrum access strategy for the secondary user communication link; the continuous long short-term memory dual duel deep Q network includes: a long short-term memory network and a dual duel deep Q network connected in sequence; the dual duel deep Q network introduces continuous learning and is constructed based on a dual deep Q network and a duel deep Q network.
9. A computer device comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the dynamic spectrum access method according to any one of claim 7.
10. A non-volatile computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the dynamic spectrum access method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Cognitive wireless network throughput optimization algorithm based on optical frequency band selection
CN107580327A
Dynamic spectrum access method based on bidirectional long-short-term memory network
CN112672359A
Dynamic spectrum access method based on deep multi-user DRQN
CN113891327A
Cognitive radio network dynamic spectrum access method based on deep reinforcement learning
CN115190489A
Advantage strategy-based vehicle networking resource allocation joint optimization method
CN118828699A
Cited By
Industrial network working condition coupling frequency spectrum avoiding method and system
CN120583529A
Vehicle-mounted radio frequency band management method, system, equipment and medium
CN120835279A