CVN Spectrum Scheduling Method and System Based on Driving State Priority and Scenario Simulation

By adopting a spectrum scheduling method based on driving state priority and scene simulation in CVN, the Monte Carlo search tree algorithm is used to build the optimal spectrum allocation solution, which solves the problem of low spectrum resource utilization on the base station side and realizes efficient spectrum resource allocation and optimization.

CN115278693BActive Publication Date: 2025-06-10DONGHUA UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210862554.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-21
Publication Date
2025-06-10
Estimated Expiration
2042-07-21

AI Technical Summary

Technical Problem

In the prior art, the base station side spectrum resource utilization rate is low, and the existing spectrum allocation method fails to effectively utilize the vehicle's driving state information and the characteristics of the network scenario, resulting in insufficient reliability and efficiency of the spectrum allocation scheme.

Method used

The CVN spectrum scheduling method based on driving state priority and scene simulation is adopted. By calculating the priority service order list of cognitive vehicles, and using the Markov decision-making process to build a Monte Carlo search tree algorithm framework, the tree strategy, simulation based on differentiated scenarios and backpropagation processes are successively implemented to obtain the optimal spectrum allocation scheme for cognitive vehicle networking.

Benefits of technology

Adaptive learning of spectrum scheduling schemes in unknown network traffic environments is realized, and approximate optimal solutions are quickly given, which greatly improves the link capacity and communication quality of cognitive vehicle users in the cellular network and improves the utilization rate of spectrum resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115278693B_ABST
    Figure CN115278693B_ABST
Patent Text Reader

Abstract

A CVN spectrum scheduling method and system based on driving state priority and scenario simulation provided by the present invention include the steps of: calculating a priority service order list of the cognitive vehicles according to the driving states and geographical dispersion degrees of the cognitive vehicles; constructing a Monte Carlo search tree algorithm framework using a Markov decision process based on the priority service order list; and sequentially iteratively executing a tree policy, simulation based on a differentiated scenario, and a backpropagation process using the Monte Carlo search tree algorithm to obtain an optimal spectrum allocation scheme for the cognitive vehicle network. The cognitive vehicle network spectrum scheduling method based on driving state priority and scenario simulation provided by the present invention can achieve adaptive learning of a spectrum scheduling scheme in an unknown network traffic environment, quickly give an approximate optimal solution, greatly improve the link capacity and communication quality of cognitive vehicle users in a cellular network, and improve the utilization rate of spectrum resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a spectrum scheduling method and system for CVN (Cognitive Vehicular Networks) based on driving state priority and scenario simulation, belonging to the technical field of vehicular networks. Background Art

[0002] At present, vehicular networks are regarded as an emerging concept in intelligent transportation network systems. Briefly speaking, it is an information network for communication between vehicles or between vehicles and infrastructure. With the rapid development of ubiquitous networks and intelligent transportation systems, the efficient communication of the Internet of Everything has attracted extensive attention. The rapid development of intelligent transportation has put forward higher requirements for the security services and entertainment services of vehicular networks, resulting in an explosive growth in communication demands of vehicular networks, which has caused a series of problems.

[0003] Firstly, the problems of insufficient utilization of authorized spectrum resources and shortage of spectrum resources for vehicular networks are important reasons restricting the development and implementation of vehicular network technologies. To meet the spectrum resource requirements of large-scale urban vehicular networks in the future, a more efficient dynamic spectrum resource scheduling scheme is needed. Relying on cognitive vehicular network technologies, we attempt to reasonably allocate the idle spectrum resources under cellular networks to cognitive vehicle users.

[0004] Regarding the dynamic spectrum allocation in mobile wireless networks, there have been many related studies. Among them, the mainstream research methods can be mainly divided into four categories: (1) allocation methods based on traditional optimization theories; (2) allocation methods based on game theory; (3) particle swarm algorithms based on swarm intelligence optimization; (4) allocation methods based on machine learning. Although the above methods can be used to solve the spectrum allocation problem, there are also many obvious deficiencies. Firstly, when the constraint conditions are complex, traditional optimization theories and game theory are not suitable for quickly solving large-scale dynamic programming problems. Secondly, swarm intelligence optimization is prone to falling into local optima. In addition, the setting and selection of effective parameters in swarm intelligence optimization algorithms are also complex. Recently, the Deep Reinforcement Learning (DRL) algorithm has been proven to be able to solve complex dynamic decision-making problems with high-dimensional state and action spaces. With the idea of the trial-and-error method, it can understand the potential laws in the environment, thus assisting intelligent decision-making. However, this machine learning-based method also has some limitations, such as slow learning speed, poor convergence, and poor adaptive ability.

[0005] In addition, consider the following two aspects: on the one hand, the above work does not incorporate the driving state information of vehicles into the process of spectrum resource scheduling, and incorporating the vehicle environmental background information can improve the reliability of the spectrum allocation scheme. On the other hand, existing spectrum allocation methods often rely on expert experience knowledge and complex parameter settings, and the portability of the search structure is not strong, so a more efficient search structure is needed. Existing research work on cognitive radio spectrum allocation does not evaluate the mobility of vehicle users themselves and the additional benefits of network scenarios and incorporate them into the solution process. Summary of the Invention

[0006] The object of the present invention is to solve the problem of low utilization rate of spectrum resources on the base station side in the prior art.

[0007] To achieve the above object, a technical solution of the present invention is to provide a CVN spectrum scheduling method based on driving state priority and scenario simulation, which is characterized by including the following steps:

[0008] S1: Calculate a priority service order list of cognitive vehicles according to the driving state and geographical dispersion degree of cognitive vehicles;

[0009] S2: Based on the priority service order list, use the Markov decision process to construct a Monte Carlo search tree algorithm framework, including the following steps:

[0010] Define the state space and action space of the Markov decision process according to the following formula:

[0011]

[0012] In the formula, s v represents the state value of node v, which is composed of λ v , ξ v ; represents the remaining bandwidth vector on the base station side, represents the remaining bandwidth of channel m; represents the number of cognitive vehicles requesting to be allocated; ξ v represents the total bandwidth requirement of m cognitive vehicles; the action a

[0013] represents that the agent allocates channel m to a vehicle that can enter the allocation sequence currently; M represents the total number of channels;

[0014] Perform spectrum allocation for cognitive vehicles in sequence according to the priority service order list, expand child nodes, and update node state values to form a Monte Carlo search tree algorithm framework;

[0015] S3: Use the Monte Carlo search tree algorithm to iteratively execute the tree policy, simulation based on different scenarios, and backpropagation processes in sequence to obtain the optimal spectrum allocation scheme for the cognitive vehicle network. Among them, the tree policy includes selection and constraint-oriented expansion, which specifically includes the following steps:

[0016] When performing the selection process, start from the root node. When it is necessary to select which child node the current node will descend to, recursively select child nodes using the Upper Confidence Bound for Trees (UCT) of the Monte Carlo search tree. Finally, regard the child node with the largest UCT as the current node for the next expansion;

[0017] When the selection process reaches termination, perform the constraint-oriented expansion operation:

[0018] Judge whether the access count of the current node is 0. If the access count then directly enter the simulation stage; if the access count Enumerate all available actions. When enumerating, prune the action space according to the constraint conditions defined by the following formula to obtain all available actions from the current node:

[0019]

[0020] In the formula: K represents the total number of primary users k; cognitive vehicle n is a secondary user, and N is the total number of secondary users; M represents the total number of channels m; the channel availability matrix L = {l n,m |l n,m ∈ {0, 1}} N×M , when channel m is available for secondary user n, l n,m = 1; conversely, when channel m is not available for secondary user n, l n,m = 0; the secondary user interference matrix C = {c n,n′,m |c n,n′,m ∈ {0, 1}} N×N×M , c n,n′,m = 1 indicates that there is mutual interference when secondary users n and n' share channel m for information transmission, and c n,n′,m = 0 indicates that secondary users n and n' can use channel m simultaneously under the condition of meeting the interference-free constraint; the channel allocation matrix A = {a n,m |a n,m ∈ {0, 1}} N×M , a n,m = 1 indicates that channel m is allocated to secondary user n, a n,m= 0 is regarded as not allocating channel m to secondary user n; the channel reward matrix R = {r n,m | r n,m ≥ 0} N×M where r n,m represents the network reward obtained by secondary user n when using channel m; P m,k,n represents the interference power of secondary user n received by primary user k on channel m; δ m,k represents the maximum acceptable interference power of primary user k on channel m; U(A, R) represents the total link capacity of the network system, where A m , R m represent the m-th column vectors of the channel allocation matrix A and the channel reward matrix R respectively, and the operation symbol represents the Hadamard product, and SUM is the operator that returns the sum of all entries in the matrix; represents the transmission power of secondary user n on channel m, and represent the minimum and maximum allowable transmission powers of secondary user n on channel m respectively; φ m represents the available bandwidth threshold of channel m, represents the transposed vector of R m ;

[0021] Then, new nodes are added to expand the Monte Carlo search tree, and the current node is set to a newly randomly selected child node after expansion;

[0022] If the number of visits to the current node is 0, then a simulation from the current node to the terminal leaf node is performed, and the current node is the newly expanded node The terminal leaf node is represented by , so when simulating, the network service duration τ of the primary user is incorporated into the reward evaluation of multi-stage expansion in the simulation process. Let the service duration τ k of primary user k correspond to an uncertainty scenario π k , and the network service duration of the primary user follows a lognormal distribution; χ samplings are performed at each layer of simulation to control the computational scale, and a set of scenarios is obtained, represented as Then the simulation based on differentiated scenarios includes the following steps:

[0023] When allocating channel m to cognitive vehicle n, the search tree performs a simulation from node to the next node, and at this time, the random revenue of node is:

[0024]

[0025] where: E represents the expectation of the random revenue obtained by cognitive vehicle n in χ scenarios; τ iis one of the samples from the distribution, 1 ≤ i ≤ χ, τ i -1 characterizes the relationship between the primary user service duration and the vehicle user revenue; utility n > 0 is a weight coefficient representing the network utility score of cognitive vehicle n. The value of utility of cognitive vehicle n is normalized to the interval [0, 1] using the hyperbolic tangent function tanh(·); Count(L n ) records the number of elements equal to 1 in the m-th column of the channel availability matrix L, Count(A m ) records the number of elements equal to 1 in the m-th column of the channel allocation matrix A, Count(L m ) - Count(A m ) describes the maximum number of vehicle users that can access channel m without considering the interference constraint C and the capacity constraint φ m . λ m represents the remaining bandwidth of channel m, m measures the remaining minimum average bandwidth that cognitive vehicle n can currently obtain; measures the remaining minimum average bandwidth that cognitive vehicle n can currently obtain;

[0026] During the simulation phase, the reward Q is adjusted for the node : v′ :

[0027]

[0028] where r n,m is the immediate reward for allocating channel m to cognitive vehicle n;

[0029] When the simulation reaches the terminal leaf node , the cumulative simulation reward of all nodes on the simulation path from node to the terminal leaf node is obtained, that is: That is:

[0030]

[0031] When an iteration reaches the terminal leaf node , the cumulative simulation reward is obtained and backpropagation is performed. The purpose of backpropagation is to update the empirical information of the prior exploration of the search tree before the next iteration. The reward of backpropagation includes the reward evaluation of the expanded nodes on all simulation paths, reflecting the overall spectrum allocation performance of the simulation strategy in the current iteration;

[0032] After reaching the iteration termination condition, the optimal spectrum allocation scheme of the current cognitive vehicular network is output.

[0033] Preferably, the step S1 includes the following steps:

[0034] Step S11: For a cognitive vehicle n that initiates a service request, calculate the vehicle travel evaluation score Travelingscore based on its driving direction, GPS coordinates, speed, and acceleration n :

[0035]

[0036] where θ n is the angle between the line connecting the GPS coordinate position of the cognitive vehicle n and the base station position and the current driving direction of the vehicle; v n represents the speed of the cognitive vehicle n, and v min , v max represent the minimum and maximum values of the driving speed of the cognitive vehicle n respectively; a n represents the acceleration of the cognitive vehicle n;

[0037] S12: Calculate the network utility score Utility of the vehicle according to the geographical dispersion degree of the cognitive vehicle n n :

[0038]

[0039] where SNR n is the signal-to-noise ratio of the signal received by the receiver of the cognitive vehicle n from the base station; log 2 (1 + SNR n ) represents the data reception rate of the vehicle n in the limited bandwidth, that is, the throughput that the vehicle can achieve; Dispersion n,n′ represents the dispersion degree between the cognitive vehicle n and the cognitive vehicle n′, and ∑ 1≤n,n′≤N,n≠n′ Dispersion n,n′ represents the global user dispersion degree of the cognitive vehicle n within the coverage area of the base station.

[0040] S13: Calculate the comprehensive priority evaluation score priorityscore of the vehicle according to the vehicle travel evaluation score and the network utility score n :

[0041] priorityscore n = Travelingscore n · Utility n

[0042] S14: Sort the comprehensive priority evaluation scores of different vehicles from large to small to obtain the priority service order list of cognitive vehicles in the current allocation period.

[0043] Preferably, in step S12, the dispersion between cognitive vehicle n and cognitive vehicle n′ n,n′ is defined as:

[0044]

[0045] where ε n represents the dispersion threshold; D n,n′ represents the average dispersion time between cognitive vehicle n and cognitive vehicle n′.

[0046] Preferably, in step S12, the average dispersion time D between cognitive vehicle n and cognitive vehicle n′ n,n′ is defined as:

[0047]

[0048] where β n,n′ (t) represents the communication dispersion state between cognitive vehicle n and cognitive vehicle n′: when there is communication interference between cognitive vehicle n and cognitive vehicle n′ in terms of geographical location, β n,n′ (t) = 0, indicating that they are in a meeting state; when there is no communication interference between cognitive vehicle n and cognitive vehicle n′ in terms of geographical location, then γ n,n′ (t) = 1, indicating that they are in a dispersed state; represents the total dispersion time between cognitive vehicle n and cognitive vehicle n′ within an allocation period T; τ n,n′ represents the statistical number of times that cognitive vehicle n and cognitive vehicle n′ are in a dispersed state within an allocation period T.

[0049] Preferably, in step S2, the cognitive vehicles are spectrally allocated in sequence according to the priority service order list, and the sub - nodes are expanded and the node state values are updated. The formation of the Monte Carlo search tree algorithm framework includes the following steps:

[0050] Create the root node v of the Monte Carlo search tree and initialize the node state value of the root node where is the number of times the node v is visited, s v is the environmental state value, and Q v is the cumulative reward value obtained by the node v;

[0051] Starting from the root node v, the spectrum of each cognitive vehicle is allocated in sequence according to the priority service order list. Each layer expansion of the Monte Carlo search tree represents the spectrum allocation of a cognitive vehicle; when the channel allocation action of the current cognitive vehicle is performed, the Monte Carlo search tree expands downward to the sub - node and updates the node state value of the sub - node until the tree expansion reaches the iteration termination condition, and the iteration terminates;

[0052] When the Monte Carlo search tree expands from one node to the next, an offline environment state predictor based on a deep neural network is used to obtain the environment state value s of the current node v v and the channel allocation action a of the current cognitive vehicle m to obtain the predicted value of the environment state of the next node v′ Then there is:

[0053]

[0054] In the formula, w ESP is the parameter of the deep neural network, and f ESP is the state-action transition function

[0055] Preferably, when training the offline environment state predictor, during a period of time after the cold start stage of the Monte Carlo search tree algorithm, the state-action transition pairs obtained through the base station are used as training data and input into the offline environment state predictor to obtain the state-action transition function f ESP .

[0056] Preferably, in step S3, the selection criterion for the optimal child node is:

[0057]

[0058] In the formula, c≥0 is a coefficient used to adjust the exploration and exploitation weights; child(v) represents the set of child nodes in the Monte Carlo search tree with the current node v as the parent node; respectively represent the total number of times the child node v′ and its parent node v are iteratively visited; Q v ′ represents the cumulative reward obtained by the child node v′

[0059] Preferably, use to represent the set of optional actions for starting the next round of channel allocation for the current cognitive vehicle n to be allocated from node v, that is, the interference-free action space of the current cognitive vehicle n to be allocated. Then, in step S3, when pruning the action space, the following steps are used for action pruning:

[0060] Introduce the channel availability matrix L of the current cognitive vehicle n to be allocated into the search tree for pruning to reduce the set of optional actions, that is, map the elements with l n,m =1 in the channel availability matrix L of the cognitive vehicle n to the set of optional actions;

[0061] Introduce the secondary user interference matrix C between cognitive vehicles into the tree search for pruning the tree structure, and judge whether a n′,m =1 and c n,n′,m =1 hold simultaneously. If both conditions hold, then the channel allocation action a in the action spacem Set of removal actions;

[0062] Judge whether the constraint conditions defined in step S3 are all satisfied. If the available channel m of the currently to-be-allocated cognitive vehicle n does not meet these constraints, then remove the channel allocation action a m from the set of optional actions;

[0063] If skip the current allocation and wait for the next round of allocation.

[0064] Preferably, in step S3, the node state values on the path from the root node to the expansion node are updated according to the following statistical rules:

[0065] Another technical solution of the present invention is to provide a CVN spectrum scheduling system based on driving state priority and scenario simulation, which is characterized in that it is used to implement the above-mentioned CVN spectrum scheduling method based on driving state priority and scenario simulation, including:

[0066] Priority calculation module, which is used to calculate the priority service order list of the cognitive vehicle according to the driving state and geographical dispersion degree of the cognitive vehicle;

[0067] Algorithm construction module, which is used to construct the Monte Carlo search tree algorithm framework using the Markov decision process according to the priority service order list;

[0068] Algorithm execution module, which is used to sequentially iterate and execute the tree policy, simulation based on different scenarios, and backpropagation process using the Monte Carlo search tree algorithm to obtain the optimal spectrum allocation scheme of the cognitive vehicle network.

[0069] A cognitive vehicle network spectrum scheduling method and system based on driving state priority and scenario simulation provided by the present invention have the following effects: it can realize the adaptive learning of the spectrum scheduling scheme in an unknown network traffic environment, quickly give an approximate optimal solution, and greatly improve the link capacity and communication quality of cognitive vehicle users in the cellular network. Brief Description of the Drawings

[0070] Figure 1 It shows a schematic diagram of the system scenario of cognitive vehicle network spectrum scheduling in an embodiment of the present invention.

[0071] Figure 2 It shows a schematic flowchart of a cognitive vehicle network spectrum scheduling method disclosed in an embodiment of the present invention.

[0072] Figure 3 It shows a schematic diagram of the search steps of Finder-MCTS in an embodiment of the present invention.

[0073] Figure 4 It shows a schematic diagram of the iterative calculation process of Finder-MCTS in an embodiment of the present invention.

[0074] Figure 5 It shows a schematic diagram of the iterative calculation process of Finder-MCTS in an embodiment of the present invention.

[0075] Figure 6 It shows a schematic diagram of the backpropagation process of Finder-MCTS in an embodiment of the present invention.

[0076] Figure 7 It shows a schematic diagram of the principle structure of the cognitive vehicular network spectrum scheduling system in the present invention. Detailed implementation manners

[0077] The present invention will be further described below in conjunction with specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. In addition, it should be understood that after reading the content taught by the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the appended claims of this application.

[0078] As Figure 1 shown, in an embodiment of the present invention, the designed allocation algorithm is deployed on the base station, and vehicle nodes equipped with cognitive radio modules can sense whether there are available spectrum resources. Vehicles can use the common control channel to send requests to access the channel to the base station. After the base station centrally collects requests from each vehicle, it learns an approximate optimal strategy for allocating available channel resources to cognitive vehicles within the coverage area. Since the vehicular network is a dynamic network, the designed spectrum resource allocation algorithm must be executed within the defined allocation time window. A large time window cannot meet the real-time requirements of the vehicular network, but a too small time window cannot support the good operation of the algorithm. In this embodiment, the size of the time window is 10s.

[0079] In the system scenario of cognitive vehicular network spectrum scheduling, the primary users are authorized mobile phone users in the current network, and the secondary users are vehicles equipped with cognitive radio modules. In this embodiment, at time t, there are N secondary users in the cognitive vehicular network competing for M channel resources, and the channels are orthogonal and non-overlapping. When the primary user k (1 ≤ k ≤ K) occupies an authorized channel m (1 ≤ m ≤ M), where K represents the total number of primary users, a protection area will be generated around the primary user, and its protection radius is Similarly, since each secondary user n (1 ≤ n ≤ N) has different transmission powers on different channels m, different interference radii will also be generated Therefore, the geographical distance between the primary user k and the secondary user n is the Euclidean distance between their coordinates, that is When the inequality constraint holds, that is, when the communication protection range generated by the primary user k on channel m intersects with the interference range generated by the secondary user n in space, it is considered that there is communication interference between the primary user k and the secondary user n; conversely, if the inequality constraint does not hold, it is considered that there is no communication interference between the primary user k and the secondary user n. Similarly, the Euclidean distance between two different secondary users n and n' is When the inequality holds, there is communication interference between two different secondary users n and n'; conversely, it is considered that there is no communication interference between two different secondary users n and n'. Communication interference affects the spectrum allocation principle of multi-channel multi-users, that is, while the primary user has the highest priority usage right of the channel, when there is no communication interference between two users, the two can use the same channel for channel transmission; conversely, the two cannot access and use the same channel simultaneously.

[0080] Secondly, the spectrum resource allocation model consists of a channel availability matrix L, a secondary user interference matrix C, a conflict-free channel allocation matrix A, and a channel reward matrix R. First, the channel availability matrix L = {l n,m |l n,m ∈{0,1}} N×M is a matrix used to represent the channel availability. When channel m is available to secondary user n, l n,m = 1; conversely, when channel m is not available to secondary user n, l n,m = 0. And in order to determine whether channel m is available to secondary user n, two aspects of analysis are required: First, when the inequality holds, secondary user n cannot use the channel occupied by primary user k; secondly, secondary user n needs to compare the interference power it receives on channel m with the interference power threshold at which it can achieve effective transmission. If the following inequality is satisfied, it is considered that channel m is available to secondary user n:

[0081]

[0082] In formula (1), P m,n,k is the power of the primary user k received by secondary user n from channel m; N m represents the environmental background noise level of channel m; γ m represents the maximum acceptable interference level on channel m.

[0083] The secondary user interference matrix C = {c n,n',m |c n,n',m ∈{0,1}} N×N×M is a matrix used to describe the interference between two different secondary users n and n'. c n,n',m= 1 indicates that there is mutual interference when secondary users n and n' share channel m for information transmission; conversely, c n,n',m = 0 indicates that secondary users n and n' can simultaneously use channel m under the condition of satisfying the interference-free constraint. In particular, for n = n', we have c n,n,m = 1 - l n,m holds. Moreover, the matrix elements should satisfy c n,n',m ≤ l n,m × l n',m , that is, interference can only occur when channel m is simultaneously available to secondary users n and n'.

[0084] The channel allocation matrix A = {a n,m |a n,m ∈ {0, 1}} N×M is a matrix used to describe the conflict-free channel allocation result of secondary users. Among them, a n,m = 1 means that channel m is allocated to secondary user n; conversely, a n,m = 0 is regarded as not allocating channel m to secondary user n. At the same time, the establishment of the channel allocation matrix A must also satisfy the interference constraint given by the secondary user interference matrix C, that is, for two different secondary users n and n', when c n,n',m = 1, the equation a n,m ·a n',m = 0 holds. In addition, in this embodiment, each secondary user can only occupy one channel for information transmission. Therefore, for any two different channels m and m', the decision variable of any secondary user n ∈ N should satisfy the inequality a n,m + a n,m' ≤ 1.

[0085] The channel reward matrix R = {r n,m |r n,m ≥ 0} N×M is a matrix used to represent the channel rewards of different secondary user links. r n,m represents the network reward obtained by secondary user n when using channel m, which is evaluated by the link capacity and is expressed as:

[0086] r n,m = W m ·log 2 (1 + SINR n,m ) (2)

[0087] In formula (2), W m is the bandwidth of channel m, and SINR n,m represents the link signal-to-noise ratio of secondary user n accessing channel m, and its calculation formula is as follows:

[0088]

[0089] In Equation (3), A m represents the m-th column of the channel allocation matrix A; Count(A m ) represents the total number of secondary users allocated on channel m; P m,n represents the signal power received by the receiver (base station) from the transmitter (secondary user n) on channel m.

[0090] From the above definitions and analyses, it can be seen that there is more than one channel allocation matrix A that satisfies the allocation restriction conditions. Therefore, there are also multiple corresponding spectrum allocation schemes, and different spectrum allocation schemes will result in different total system rewards. The objective of the present invention is to solve an optimal channel allocation matrix A * , such that when spectrum allocation is performed according to the allocation scheme corresponding to the optimal channel allocation matrix A * , the maximum total network system link capacity U(A, R) can be obtained. The total link capacity is defined as:

[0091]

[0092] In Equation (4), A m , R m respectively represent the m-th column vectors of the channel allocation matrix A and the channel reward matrix R; the operation symbol represents the Hadamard product, that is, the product of the corresponding position elements of two vectors; SUM is an operator that returns the sum of all entries of the matrix. A m can be regarded as an N×1-dimensional 0 / 1 decision vector, R m is an N×1-dimensional real reward vector, is an N×1-dimensional vector.

[0093] Therefore, by optimizing the quality of the allocation scheme for accessing secondary users and allocating the available spectrum resources to more reasonable cognitive users, the problem of low utilization rate of spectrum resources on the base station side can be solved. In summary, we describe this combinatorial optimization problem of spectrum resource allocation as a binary integer linear programming problem (Binary Integer Linear Programming, BILP) shown in Equation (4) and Equations (5.1)-(5.8):

[0094]

[0095] Among them, Equation (5.1) gives the input matrix vectors A m and R mThe element value range; Equation (5.2) ensures that the channel allocated to secondary user n in the allocation scheme must be an available channel; To avoid link conflicts on the specified channel m and protect the communication of each cognitive user from interference by other cognitive users, the elements of the conflict-free channel allocation matrix for this channel should satisfy Equation (5.3). Equation (5.4) indicates that each secondary user can only occupy one channel for information transmission. In the constraint condition (5.5), represents the transmission power of secondary user n on channel m, and respectively represent the minimum and maximum allowable transmission powers of secondary user n on channel m. In the constraint condition (5.6), P m,k,n represents the interference power received by primary user k from secondary user n on channel m, and δ m,k represents the maximum acceptable interference power of primary user k on channel m. For any primary user k, the total received interference signal power on the channel m it occupies must be kept below the maximum acceptable interference threshold, that is, to ensure that the primary user is not interfered by the communication of secondary users on this channel. In the constraint condition (5.7), φ m represents the available bandwidth threshold of channel m, represents the transpose vector of R m , and this constraint ensures that the total network capacity of the links accessed by channel m should be less than or equal to its available bandwidth.

[0096] As Figure 2 shown, the present invention provides a cognitive vehicle networking spectrum scheduling method based on driving state priority and scenario simulation, including the following steps:

[0097] S1: Calculate the priority service order list of cognitive vehicles according to the driving state and geographical dispersion degree of cognitive vehicles;

[0098] S2: Based on the priority service order list, construct a Monte Carlo search tree algorithm framework using the Markov decision process;

[0099] S3: Use the Monte Carlo search tree algorithm to sequentially iterate through the tree policy, simulation based on differentiated scenarios, and backpropagation process to obtain the optimal spectrum allocation scheme for the cognitive vehicle networking, where the tree policy includes selection and constraint-oriented expansion.

[0100] The cognitive vehicle networking spectrum scheduling method provided by the present invention obtains the comprehensive priority evaluation score of each vehicle by defining a vehicle driving evaluation score and a network utility score. According to the priority score, the available spectrum resources are allocated from the highest priority to the lowest vehicle users, which can improve the performance in dynamic spectrum allocation in vehicle networking. Then, combined with the priority scoring, we present a Monte Carlo search tree algorithm based on different vehicle networking scenarios. Compared with the classical Monte Carlo Tree Search (MCTS), this algorithm provides a constraint-oriented tree expansion and scenario simulation mechanism, which can achieve the adaptive learning of spectrum scheduling schemes in unknown network traffic environments and quickly give approximate optimal solutions, greatly improving the link capacity and call quality of cognitive vehicle users in cellular networks. To distinguish it from the classical MCTS, the algorithm in the present invention is called Finder-MCTS.

[0101] Further, step S1 includes:

[0102] S11: For a secondary user n that initiates a service request, calculate the vehicle driving evaluation score Travelingscore of the vehicle according to its driving direction, GPS coordinates, speed, and acceleration n :

[0103]

[0104] In formula (6), θ n is the included angle between the line connecting the GPS coordinate position of vehicle n and the base station position and the current driving direction of the vehicle; v n represents the speed of vehicle n, v min , v max respectively represent the minimum and maximum values of the driving speed of vehicle n; a n represents the acceleration of vehicle n.

[0105] S12: Calculate the network utility score Utility of the vehicle according to the geographical dispersion degree of the vehicle n :

[0106]

[0107] In formula (7), SNR n is the signal-to-noise ratio of the received signal from the base station by the receiver of vehicle n; log 2 (1 + SNR n ) represents the data reception rate of vehicle n in the limited bandwidth, that is, the throughput that the vehicle can achieve; ∑ 1≤n,n′≤N,n≠n′ Dispersion n,n' represents the global user dispersion degree of vehicle n within the coverage area of the base station.

[0108] S13: Calculate the comprehensive priority evaluation score priorityscore of the vehicle based on the vehicle driving evaluation score and the network utility score n :

[0109] priorityscore n = Travelingscore n · Utility n (8)

[0110] S14: Sort the comprehensive priority evaluation scores of different vehicles from large to small to obtain the priority service order list of cognitive vehicles in the current allocation period

[0111] Obviously, a relatively large included angle θ n indicates that vehicle n will drive out of the base station coverage area in a relatively short time in the future. Therefore, cognitive vehicles with relatively large included angles should be given relatively low spectrum allocation weights; cognitive vehicles with relatively small included angles should be given relatively high spectrum allocation weights. We use to normalize the included angle to the interval [0,1]. In addition, for vehicles with relatively high driving speeds, they will drive out of the base station coverage area in a short time in the future and should be given relatively low spectrum allocation weights; on the contrary, vehicles with relatively low driving speeds should be given relatively high spectrum allocation weights. Therefore, we adopt the normalization formula to describe the impact of the driving speed of vehicle n on the spectrum allocation weight. In addition, the magnitude of the acceleration in the same direction is positively correlated with the change in the magnitude of the speed. Therefore, for vehicles with relatively large driving accelerations, their driving speeds will increase significantly, so they should be given relatively low spectrum allocation weights; on the contrary, vehicles with relatively small accelerations should be given relatively high spectrum allocation weights. To normalize to the interval [0,1], we adopt to emphasize the impact of vehicle driving acceleration on the vehicle spectrum allocation weight. In addition, in order to constrain the value of Travelingscore n to the interval [0,1], it is also necessary to multiply by Finally, formula (6) is obtained. The higher the vehicle driving evaluation score, the more stable the future access state of the vehicle's local state in the network, and the longer the vehicle will stay in the current cognitive vehicle network. Therefore, the vehicle will obtain a higher spectrum allocation weight

[0112] Furthermore, in step S12, the dispersion between two different vehicles n and n' is defined as:

[0113]

[0114] In formula (9), εn represents the dispersion threshold; D n,n' represents the average dispersion time between two different vehicles n and n', defined as:

[0115]

[0116] In Equation (10), β n,n' (t) represents the communication dispersion state between two different vehicles n and n'. When there is communication interference between two different vehicles n and n' geographically, β n,n' (t) = 0, indicating that they are in the "encounter" state; conversely, if there is no such communication interference between two different vehicles n and n' at the geographical location, then β n,n' (t) = 1, indicating that they are in the "dispersion" state. represents the total dispersion time of two different vehicles n and n' within an allocation period T; τ n,n' represents the statistical number of times that two different vehicles n and n' are in the "dispersion" state within an allocation period T. Obviously, the average dispersion time D n,n' The larger the value, the longer the time that the secondary users n and n' are in the "dispersion" state, and then the global user dispersion degree ∑ of vehicle n 1≤n,n'≤N,n≠n' Dispersion n,n' is larger, and at this time the network utility score of the user is higher. A vehicle with a larger network utility score means relatively better global communication ability, so the vehicle will also obtain a higher spectrum allocation weight.

[0117] In this embodiment, the dispersion threshold ε n is obtained by taking the median of the average dispersion times of all vehicles within the coverage of the current base station.

[0118] In an embodiment of the present invention, the vehicle driving evaluation score and the network utility score are obtained by the base station through real-time collection and analysis of relevant characteristic information of vehicles in the network. Therefore, for vehicle users in the network that initiate service requests to the base station, the base station uses the collected vehicle information to calculate the priority scores of the requested vehicles, and sorts them from large to small, so as to obtain the priority service order list of users in the cognitive network within the current allocation period. This priority service order list will be used as the allocation order list for secondary users, thereby ensuring that vehicle users with different allocation weights in the network access according to priority, and improving the reliability of the spectrum scheduling scheme.

[0119] Furthermore, step S2 includes:

[0120] Define the state space and action space of the Markov decision process according to the following formulas respectively:

[0121]

[0122] In formula (11), s v represents the state value of node v, which consists of three parts: λ v , ξ v , where: represents the remaining bandwidth vector on the base station side, represents the remaining bandwidth of channel m; represents the number of vehicles for which requests are to be allocated; ξ v represents the total bandwidth requirement of m vehicles; the action a

[0123] Based on the state space and action space, Finder-MCTS is constructed. Among them, Finder-MCTS consists of nodes and edges. Each node maintains a node state value, which includes three types of statistical information: the number of times the node is visited, the state value, and the cumulative reward value obtained by the node; the edge represents the action that causes the state transition.

[0124] Perform spectrum allocation for vehicles in sequence according to the priority service order list, expand child nodes, and update the node state value to form the algorithm framework of Finder-MCTS.

[0125] As Figure 3 shown, the search steps of Finder-MCTS in an embodiment of the present invention include:

[0126] First, create the root node v of the search tree and initialize the node state value of the root node where, is the number of times node v is visited, s v is the state value, and Q v is the cumulative reward value obtained by node v. In this embodiment, the priority service order list of vehicles is vehicle ID 3 , vehicle ID 1 , vehicle ID 2 , and each layer expansion of the search tree represents the spectrum allocation for a vehicle. Therefore, starting from the root node v, first perform spectrum allocation for vehicle vehicle ID 3 . When the channel allocation action a 3 of vehicle vehicle ID 1 occurs, the search tree expands downward to the child node v', and the node state value of the child node v' is updated to where each allocation process includes an iterative calculation process of selection, expansion, simulation, and backpropagation; when the spectrum allocation for vehicle vehicle ID3 After the allocation, immediately start from node v' to perform spectrum allocation for vehicle ID 1 When performing spectrum allocation for vehicle ID 1 with channel allocation action a 5 , the search tree expands downward to child node v'', and updates the node state value of child node v''; and so on. When the tree expansion reaches the iteration termination condition, that is, when all secondary users are allocated or there is no remaining available bandwidth resource for channel allocation, the iteration terminates. At this time, the allocation path indicated by the black arrow line is v→v'→v''→v''', and the set of allocation strategies formed by its corresponding actions is {a 1 , a 5 , a 1}. According to the priority service order list and the set of allocation strategies, a channel allocation matrix A N×M can be obtained.

[0127] Due to the uncertainty of the spectrum occupancy activities of primary users, when the tree expands from one node to the next node, the expansion will be unstable, that is, given a state and an action, the next state is uncertain. Therefore, in order to limit the horizontal expansion scale of the search tree and speed up the search, it is necessary to gradually learn a real environment model close to the CVN when performing spectrum allocation. Thus, the present invention provides an offline environmental state predictor (ESP) based on a deep neural network (DNN). Since obtaining the ESP requires sufficient training data, in the cold start phase of Finder-MCTS, that is, at the beginning of the algorithm operation, the ESP is not used; and within a period of time after the cold start phase, the base station can calculate and obtain a considerable number of "state-action transition pairs" in real time. Subsequently, we continuously input these "state-action transition pairs" as training data into the ESP, so that the state-action transition function f ESP can be obtained, and this is an offline training process. When f ESP is available, the search of Finder-MCTS will converge faster due to the reduction of branches.

[0128] The network structure of the DNN consists of an input layer, three hidden layers, and an output layer. In this embodiment, the learning rate of the DNN is set to 0.05, the activation function of the DNN is selected as the rectified linear unit (Relu), and the Mini-Batch gradient descent method is used to optimize the network parameters of the DNN to ensure the speed and accuracy of training convergence. First, in the DNN, the training label is the real environmental state s v ', which is the state of the node v corresponding to the expanded child node v'; secondly, use the ESP to predict the state at the next moment The loss function is as follows:

[0129]

[0130] In Equation (12), B represents the size of the batch in Mini-Batch gradient descent. In this embodiment, B = 64, indicating that 64 samples are selected for each iteration; ||·|| 2 represents the L2 norm.

[0131] After the loss function converges, update the parameters w of the optimized DNN network ESP , and then based on the selected action a m and state s v , use ESP to obtain the environmental state of the expanded node:

[0132]

[0133] As Figure 4 , 5 shown, Finder-MCTS needs to sequentially iterate through the selection, expansion, simulation, and backpropagation processes, and selection and expansion are together called the tree policy.

[0134] First, when Finder-MCTS executes the selection process, starting from the root node, when the algorithm must choose which child node it will descend to, the algorithm tries to find a good balance between exploration and exploitation. In an embodiment of the present invention, we use the Upper Confidence Bound for Trees (UCT) to recursively select child nodes. Eventually, the algorithm will regard the child node with the largest UCT value as the current node for the next expansion. The selection criterion for the optimal child node is:

[0135]

[0136] In Equation (14), c ≥ 0 is a coefficient used to adjust the exploration and exploitation weights. In this embodiment, c = 0.8 is set. child(v) represents the set of child nodes in the Finder-MCTS tree with v as the parent node. respectively represent the total number of times the child node v' and its parent node v are iteratively visited. Q v ' represents the cumulative reward obtained by the leaf node v'. It should be noted that the selected child node should be expandable, that is, it has unvisited child nodes and represents a non-terminal state.

[0137] When the selection process reaches termination, the algorithm will then perform the expansion operation. The algorithm first determines whether the access count of the current node is 0. If the access count then the algorithm directly enters the simulation phase; if the access count The algorithm will enumerate all available actions. However, if simple enumeration is used, the number of optional actions at the next layer is M. Thus, as the tree expands, more action expansions will result in a huge search structure while the search tree descends, and its computational complexity grows in a geometric series relationship with the number of secondary users to be allocated in the network. Therefore, the present invention provides a constraint-oriented expansion.

[0138] In the constraint-oriented expansion, the action space is pruned according to the constraint conditions defined by formula (5) and formulas (5.1) to (5.8) so as to obtain all available actions from the current node. Then, new nodes are added to expand the tree, and the current node is set as a newly randomly selected child node after expansion.

[0139] Further, in the present invention, is used to represent the set of optional actions starting from node v for the next round of channel allocation for secondary user n, that is, the interference-free action space of secondary user n. We adopt three steps to perform action pruning. First, the availability of channels should be considered, so the available channel matrix L of the vehicle is introduced into the search tree for pruning to reduce the set of optional actions, that is, the elements with l n,m =1 in the available channel matrix of vehicle n are mapped to the set of optional actions. Second, considering that the currently to-be-allocated vehicle vehicle ID n should not share the same channel with the vehicles with communication interference, the secondary user interference matrix C between vehicles is introduced into the tree search for pruning the tree structure. At this time, the algorithm will judge whether a n′,m =1 and c n,n′,m =1 hold simultaneously. If both conditions hold simultaneously, a m in the action space will be removed from the action set. Next, the algorithm will judge whether constraints (5) and constraints (5.1) to (5.8) hold simultaneously. If the optional channel m of the currently to-be-allocated vehicle does not satisfy these constraints, a m will be removed from the set of optional actions. Finally, if the algorithm will skip the current allocation and wait for the next round of allocation.

[0140] From the above expansion process, it can be known that if the number of visits to the current node is 0, a simulation from the current node (i.e., the newly expanded node, represented by ) to the terminal leaf node (represented by ) will be executed. Generally, the simulation strategy adopts a random search strategy, and a reward is generated at the terminal leaf node However, the time-varying nature of the activities of primary users occupying the spectrum makes the actual available spectrum resources on the base station side uncertain, and this uncertainty will have a potential impact on the reward evaluation of cognitive vehicle users waiting to be assigned channels in the CVN. Therefore, in the present invention, the network service duration τ of the primary user is incorporated into the reward evaluation of the multi-stage expansion in the simulation process. Specifically, the service duration τ of primary user k k corresponds to an uncertainty scenario π k , and the network service duration of the primary user follows a lognormal distribution, that is, its probability density function is:

[0141]

[0142] In Equation (15), the parameters (μ, σ) are in milliseconds (ms). In this embodiment, the values of (μ, σ) are (2.47, 1.88).

[0143] Since there are theoretically infinitely many scenarios during sampling, in this embodiment, we perform χ samplings during each layer of simulation to control the computational scale. Therefore, a scenario set can be obtained, denoted as First, when allocating channel m to vehicle ID n , the search tree performs a simulation from node to the next node. At this time, the random reward of node is:

[0144]

[0145] In Equation (16), E represents the expectation of the random rewards obtained by vehicle ID n under χ scenarios; τ i (1 ≤ i ≤ χ) is one of the samplings from the distribution, and the larger the value of τ i , the longer the time the primary user occupies the channel, and the lower the random reward of vehicle ID n during allocation. In fact, the shorter the channel occupancy time of the primary user, the higher the uncertainty reward of the vehicle user on the same channel, and τ i -1 characterizes this relationship between the primary user service duration and the vehicle user reward. utility n > 0 is a weight coefficient representing the network utility score of the secondary user vehicle ID n , which reflects the communication ability of vehicle ID n . In this embodiment, the hyperbolic tangent function tanh(·) is used to transform the utility of vehicle ID n of nThe value is normalized to the interval [0, 1]. When the utility n is higher, the weight coefficient term is closer to 1, which indicates that vehicle users with stronger communication capabilities are more inclined to obtain higher random benefits. In addition, measures the vehicle ID n The remaining minimum average bandwidth (MHz) that can be obtained currently. Count(L m ) records the number of elements with a value of 1 in the m-th column of the channel availability matrix L, and Count(A m ) records the number of elements with a value of 1 in the m-th column of the channel allocation matrix A. Count(L m ) - Count(A m ) describes the maximum number of vehicle users that can be connected to channel m without considering the interference constraint C and the capacity constraint - m . λ m represents the remaining bandwidth of channel m.

[0146] Analyze the significance of sampling from the perspective of supply and demand. When the primary users using the same channel m have a longer service duration τ, the uncertainty of the available spectrum resources on the supply side will decrease in the long term, and the resource supply tends to be less than or equal to the demand, and the random benefit obtained by the demand side will decrease accordingly; conversely, a shorter τ means that the spectrum resource supply tends to be greater than the demand, and the random benefit obtained by the demand side will increase accordingly, and the allocation scheme will obtain high benefits at this time. Obviously, a larger random benefit reflects the possibility for vehicle users to obtain greater network benefits.

[0147] In short, if the vehicle has strong communication capabilities, the service duration of the primary user is low, and the remaining resources are sufficient, the random benefit in the cognitive vehicle network will be high.

[0148] Secondly, according to the following formula, the reward Q is adjusted for the node v' in the simulation stage:

[0149]

[0150] In Equation (16), r n,m refers to the immediate reward for allocating channel m to the vehicle ID n , which is obtained from Equation (2).

[0151] When the simulation reaches the terminal leaf node , the cumulative simulation reward of all nodes on the simulation path from node to node can be obtained, that is: That is:

[0152]

[0153] As shown Figure 6 in the figure, further, when one iteration reaches the termination node after that, the cumulative simulation reward is obtained according to formula (18) for backpropagation. The purpose of backpropagation is to update the empirical information of prior exploration of the search tree before the next iteration. In this way, the reward of backpropagation includes the reward evaluations of the expanded nodes on all simulated paths, reflecting the overall spectrum allocation performance of the simulation strategy in the current iteration. At the same time, the algorithm updates the node state values on the path from the root node to the expanded node according to the following statistical rules:

[0154]

[0155] The Finder-MCTS algorithm iteratively executes functions such as tree policy, simulation, and backpropagation to explore different spectrum allocation schemes. Finally, the algorithm outputs the optimal spectrum allocation scheme for the current cognitive vehicle network.

[0156] The step division of the above method is only for clear description. When implemented, it can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, it is within the protection scope of the present invention; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but without changing the core design of its algorithm and process, are all within the protection scope of the invention.

[0157] As shown Figure 7 in the figure, the present invention also provides a cognitive vehicle network spectrum scheduling system, including: a priority calculation module, an algorithm construction module, and an algorithm execution module. Among them, the priority calculation module is used to calculate the priority service order list of the cognitive vehicles according to the driving state and geographical dispersion degree of the cognitive vehicles; the algorithm construction module is used to construct a Monte Carlo search tree algorithm framework using the Markov decision process according to the priority service order list; the algorithm execution module is used to iteratively execute the tree policy, simulation based on differentiated scenarios, and backpropagation processes using the Monte Carlo search tree algorithm to obtain the optimal spectrum allocation scheme of the cognitive vehicle network.

[0158] It should be noted that, in order to highlight the innovative part of the present invention, modules not closely related to solving the technical problems proposed by the present invention are not introduced in this embodiment, but this does not mean that there are no other modules in this embodiment.

[0159] In addition, those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the systems described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. In the embodiments provided by the present invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation; for example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed.

[0160] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed to multiple network modules. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0161] In addition, in each embodiment of the present invention, the functional modules can be integrated into a processing module, or each module can exist physically alone, or two or more modules can be integrated into one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional units.

[0162] If the above-mentioned integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present invention. The foregoing storage medium includes various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0163] As described above, a cognitive vehicle network spectrum scheduling method and system based on driving state priority and scenario simulation provided by the present invention can achieve adaptive learning of spectrum scheduling solutions in an unknown network traffic environment, quickly give an approximate optimal solution, greatly improve the link capacity and communication quality of cognitive vehicle users in a cellular network, and improve the utilization rate of spectrum resources. Therefore, the present invention effectively overcomes various shortcomings in the prior art and has high industrial utilization value.

Claims

1. A CVN spectrum scheduling method based on driving state priority and scenario simulation, characterized in that, it includes the following steps: S1: According to the driving state and geographical dispersion degree of cognitive vehicles, calculate the priority service order list of cognitive vehicles, including the following steps: Step S11: For a cognitive vehicle n that initiates a service request, calculate the vehicle driving evaluation score Travelingscore based on its driving direction, GPS coordinates, speed, and acceleration n : Where, θ n is the included angle between the line connecting the GPS coordinate position of the cognitive vehicle n and the base station position and the current driving direction of the vehicle; v n represents the speed of the cognitive vehicle n, v min and v max respectively represent the minimum value and the maximum value of the driving speed of the cognitive vehicle n; a n represents the acceleration of the cognitive vehicle n; S12: Calculate the network utility score Utility of the vehicle based on the geographical dispersion of the cognitive vehicle n n : where, SNR n is the signal-to-noise ratio of the signal received by the receiver of cognitive vehicle n from the base station; log 2 (1 + SNR n ) represents the data reception rate of vehicle n in the limited bandwidth, that is, the throughput achievable by the vehicle; Dispersion n,n' represents the dispersion between cognitive vehicle n and cognitive vehicle n'; ∑ 1≤n,n'≤N,n≠n' Dispersion n,n' represents the global user dispersion of cognitive vehicle n within the coverage area of the base station, where the dispersion Dispersion n,n' between cognitive vehicle n and cognitive vehicle n' is defined as: Where ε n represents the dispersion threshold; D n,n' represents the average dispersion time between the cognitive vehicle n and the cognitive vehicle n'; S13: Calculate the comprehensive priority evaluation score priorityscore of the vehicle based on the vehicle driving evaluation score and the network utility score n : priorityscore n = Travelingscore n ·Utility n S14: Sort the comprehensive priority evaluation scores of different vehicles from large to small to obtain the priority service order list of cognitive vehicles within the current allocation period; S2: Based on the priority service order list, use the Markov decision process to construct a Monte Carlo search tree algorithm framework, including the following steps: Define the state space and action space of the Markov decision process according to the following formula: where s v represents the state value of node v, which is composed of λ v , ξ v ; represents the remaining bandwidth vector on the base station side, represents the remaining bandwidth of channel m; represents the number of cognitive vehicles for which requests are to be allocated; ξ v represents the total bandwidth requirement of cognitive vehicles; action a m represents that the agent allocates channel m to a vehicle that can currently enter the allocation sequence; M represents the total number of channels; Based on the state space and action space, construct a Monte Carlo search tree, which consists of nodes and edges: each node maintains a node state value, including the number of times the node is visited, the environmental state value, and the cumulative reward value obtained by the node; the edge represents the action that causes the state transition; Allocate spectra to cognitive vehicles in sequence according to the priority service order list, expand sub-nodes, and update node state values to form a Monte Carlo search tree algorithm framework; S3: Use the Monte Carlo search tree algorithm to iteratively execute the tree policy, simulation based on differentiated scenarios, and backpropagation process to obtain the optimal spectrum allocation scheme of the cognitive vehicle network. Among them, the tree policy includes selection and constraint-oriented expansion, specifically including the following steps: When executing the selection process, start from the root node. When it is necessary to select which sub-node the current node will descend to, use the Upper Confidence Bound for Trees (UCT) of the Monte Carlo search tree to recursively select sub-nodes. Finally, regard the sub-node with the largest UCT as the current node for the next expansion; When the selection process reaches termination, execute the constraint-oriented expansion operation: Determine whether the access count of the current node is 0. If the access count is 0, directly enter the simulation phase; if the access count enumerate all available actions. When enumerating, prune the action space according to the constraint conditions defined by the following formula to obtain all available actions from the current node: Where: K represents the total number of primary users k; cognitive vehicle n is a secondary user, and N is the total number of secondary users; M represents the total number of channels m; the channel availability matrix L = {l n,m |l n,m ∈ {0, 1}} N×M , when channel m is available to secondary user n, l n,m = 1; conversely, when channel m is not available to secondary user n, l n,m = 0; the secondary user interference matrix C = {c n,n',m |c n,n',m ∈ {0, 1}} N×N×M , c n,n',m = 1 indicates that there is mutual interference when secondary users n and n' share channel m for information transmission, and c n,n',m = 0 indicates that secondary users n and n' can simultaneously use channel m under the condition of satisfying the interference-free constraint; the channel allocation matrix A = {a n,m |a n,m ∈ {0, 1}} N×M , a n,m = 1 means that channel m is allocated to secondary user n, and a n,m = 0 is regarded as not allocating channel m to secondary user n; the channel reward matrix R = {r n,m |r n,m ≥ 0} N×M , r n,m represents the network reward obtained by secondary user n when using channel m; P m,k,n represents the interference power of secondary user n received by primary user k on channel m; δ m,k represents the maximum acceptable interference power of primary user k on channel m; U(A, R) represents the total link capacity of the network system, A m , R m respectively represent the m-th column vectors of the channel allocation matrix A and the channel reward matrix R, and the operation symbol represents the Hadamard product, and SUM is the operator that returns the sum of all entries in the matrix; represents the transmission power of secondary user n on channel m, and respectively represent the minimum and maximum allowable transmission powers of secondary user n on channel m; φ m represents the available bandwidth threshold of channel m, represents the transposed vector of R m ; Then, add a new node to expand the Monte Carlo search tree, and set the current node to a newly randomly selected sub-node after expansion; If the access count of the current node is 0, perform a simulation from the current node to the terminal leaf node, where the current node is the newly expanded node The terminal leaf node is represented by . During the simulation, the network service duration τ of the primary user is incorporated into the reward evaluation of multi-stage expansion in the simulation process. Let the service duration τ of the primary user k k correspond to an uncertainty scenario π k , and the network service duration of the primary user follows a lognormal distribution. χ samplings are performed at each layer of the simulation to control the computational scale, resulting in a set of scenarios, denoted as Then, the simulation based on the differentiated scenarios includes the following steps: When allocating channel m to cognitive vehicle n, the search tree performs a simulation from node to the next node, where the random payoff of node is: where: E represents the expectation of the random revenue obtained by the cognitive vehicle n in χ scenarios; τ i is one of the samples from the distribution, 1 ≤ i ≤ χ, τ i -1 characterizes the relationship between the primary user service duration and the vehicle user revenue; utility n > 0 is a weight coefficient representing the network utility score of the cognitive vehicle n. The hyperbolic tangent function tanh(·) is used to normalize the utility value of the cognitive vehicle n n to the interval [0,1]; Count(L m ) records the number of elements equal to 1 in the m-th column of the channel availability matrix L, Count(A m ) records the number of elements equal to 1 in the m-th column of the channel allocation matrix A, Count(L m ) - Count(A m ) describes the maximum number of vehicle users that can be connected to channel m without considering the interference constraint C and the capacity constraint φ m ; λ m represents the remaining bandwidth of channel m, measures the remaining minimum average bandwidth that the cognitive vehicle n can currently obtain; In the simulation phase for the nodes The reward Q is adjusted v' : where r n,m denotes the immediate reward for allocating channel m to cognitive vehicle n; When the simulation reaches the terminal leaf node the cumulative simulation rewards of all nodes on the simulation path from the node to the terminal leaf node are obtained That is When an iteration reaches the terminal leaf node the cumulative simulation reward is obtained Backpropagation is performed. The purpose of backpropagation is to update the experience information of prior exploration of the search tree before the next iteration. The reward of backpropagation includes the reward evaluations of the expanded nodes on all simulation paths, reflecting the overall spectrum allocation performance of the simulation strategy in the current iteration; After reaching the iteration termination condition, output the optimal spectrum allocation scheme of the current cognitive vehicle network.

2. The CVN spectrum scheduling method based on driving state priority and scenario simulation as described in claim 1, characterized in that, In step S12, the average dispersion time D between the cognitive vehicle n and the cognitive vehicle n' n,n' is defined as: where β n,n' (t) represents the communication dispersion state between cognitive vehicle n and cognitive vehicle n': when there is communication interference between cognitive vehicle n and cognitive vehicle n' geographically, β n,n' (t) = 0, indicating that they are in a meeting state; when there is no communication interference between cognitive vehicle n and cognitive vehicle n' geographically, then β n,n' (t) = 1, indicating that they are in a dispersed state; represents the total dispersion time of cognitive vehicle n and cognitive vehicle n' within an allocation period T; τ n,n' represents the statistical number of times that cognitive vehicle n and cognitive vehicle n' are in a dispersed state within an allocation period T.

3. The CVN spectrum scheduling method based on driving state priority and scenario simulation as described in claim 1, characterized in that, In the step S2, allocating spectra to cognitive vehicles in sequence according to the priority service order list, expanding sub-nodes, and updating node state values to form a Monte Carlo search tree algorithm framework includes the following steps: Create the root node v of the Monte Carlo search tree and initialize the node state value of the root node Among them, is the number of times node v is visited, s v is the environmental state value, Q v is the cumulative reward value obtained by node v; Starting from the root node v, allocate spectra to each cognitive vehicle in sequence according to the priority service order list. Each layer expansion of the Monte Carlo search tree represents the allocation of spectra to a cognitive vehicle; when the channel allocation action of the current cognitive vehicle is performed, the Monte Carlo search tree expands downward to a sub-node, and updates the node state value of the sub-node until the tree expansion reaches the iteration termination condition, and the iteration terminates; When the Monte Carlo search tree expands from one node to the next, an offline environment state predictor based on a deep neural network is used to obtain the environment state value s of the current node v v and the channel allocation action a of the current cognitive vehicle m to obtain the predicted environment state value of the next node v' Then we have: where, w ESP is the parameter of the deep neural network, and f ESP is the state-action transition function.

4. The CVN spectrum scheduling method based on driving state priority and scenario simulation as described in claim 3, characterized in that, When training the offline environment state predictor, the state-action transition pairs obtained through the base station within a period of time after the cold start phase of the Monte Carlo search tree algorithm are used as training data and input into the offline environment state predictor to obtain the state-action transition function f ESP .

5. A CVN spectrum scheduling method based on driving state priority and scenario simulation as claimed in claim 1, characterized in that, in step S3, the selection criterion for the optimal sub-node is: Wherein, c≥0 is a coefficient used to adjust the exploration and exploitation weights; child(v) represents the set of child nodes in the Monte Carlo search tree with the current node v as the parent node; respectively represent the total number of times the child node v' and its parent node v are iteratively visited; Q v' represents the cumulative reward obtained by the child node v'.

6. A CVN spectrum scheduling method based on driving state priority and scenario simulation as claimed in claim 1, characterized in that, Usage denotes the set of optional actions for starting the next round of channel allocation for the currently to-be-allocated cognitive vehicle n from node v, that is, the interference-free action space of the currently to-be-allocated cognitive vehicle n. Then, in step S3, when pruning the action space, the following steps are adopted for action pruning: Introduce the channel availability matrix L of the current cognitive vehicle n to be allocated into the search tree for pruning to reduce the set of optional actions, that is, map the elements with l n,m = 1 in the channel availability matrix L of the cognitive vehicle n to the set of optional actions; Introduce the secondary user interference matrix C between cognitive vehicles into the tree search for pruning the tree structure, and judge whether a n',m = 1 and c n,n',m = 1 hold simultaneously. If both conditions hold, then remove the channel allocation action a m from the action set; Determine whether the constraint conditions defined in step S3 are all satisfied simultaneously. If the optional channel m of the currently to-be-allocated cognitive vehicle n does not meet these constraints, then remove the channel allocation action a m from the set of optional actions; If the current allocation will be skipped and wait for the next round of allocation.

7. A CVN spectrum scheduling method based on driving state priority and scenario simulation as claimed in claim 1, characterized in that, In step S3, the node state values on the path from the root node to the expansion node are updated according to the following statistical rules:

8. A CVN spectrum scheduling system based on driving state priority and scenario simulation, characterized in that, for implementing the CVN spectrum scheduling method based on driving state priority and scenario simulation as claimed in claim 1, including: a priority calculation module, configured to calculate a priority service order list of the cognitive vehicle according to the driving state and geographical dispersion degree of the cognitive vehicle; an algorithm construction module, configured to construct a Monte Carlo search tree algorithm framework using a Markov decision process according to the priority service order list; an algorithm execution module, configured to sequentially execute a tree policy, a simulation based on a differentiated scenario, and a backpropagation process using the Monte Carlo search tree algorithm to obtain an optimal spectrum allocation scheme for the cognitive vehicle network.

Citation Information

Patent Citations

  • Intersection Internet of Vehicles cognitive spectrum allocation mechanism based on clustering structure

    CN110557211A

  • Internet of Vehicles cognitive spectrum allocation method based on supply and demand balance

    CN111614420A