Knowledge-driven resource scheduling method in Internet of Vehicles under the integrated space-ground scenario
By adopting a knowledge-driven resource scheduling method in the integrated air-space and earth network, and using asynchronous advantages actor and critic algorithm and space-time correlation knowledge to optimize access channel selection, the problem of insufficient timeliness and realistic feasibility of resource scheduling in the existing technology is solved, and efficient and low-cost vehicle resource scheduling is achieved.
Patent Information
- Application Number
- CN202310064834.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-01-13
AI Technical Summary
The existing wireless network communication resource allocation algorithm is difficult to effectively and uniformly manage massive communication equipment and tasks in complex communication scenarios, resulting in insufficient timeliness and practical feasibility of resource scheduling and high cost.
A knowledge-driven resource scheduling method is proposed. By obtaining the application scenarios of the integrated air-space and earth network, an optimization model is established to minimize the total delay of the vehicle's business transmission, and async advantage actor and critic algorithm combines space-time correlation knowledge and mathematical knowledge to optimize access channel selection and adjust the subnet learning rate to accelerate convergence.
It improves the network's resource scheduling efficiency and interpretability, reduces the convergence delay during vehicle access, enhances the generalization ability of neural networks, and meets the high latency and low cost requirements of integrated aerospace and earth network communication.
Smart Images

Figure CN116133127B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of vehicle networks, and specifically relates to a knowledge-driven resource scheduling method in a vehicle network in an air-ground-space integrated scenario. Background Art
[0002] With the continuous development of the Internet of Vehicles, more complex communication scenarios are constantly being proposed. The demand for seamless coverage of new in-vehicle communication services has become more urgent, and the requirements for latency are also getting higher and higher. At the same time, the service transmission process of the vehicle network consumes a lot of network resources, and network resources are limited. Therefore, it is necessary to schedule network resources efficiently. However, traditional model-driven and data-driven methods cannot effectively unify resource allocation and management for massive communication devices and tasks in the case of heterogeneous and complex networks.
[0003] In the existing wireless network communication resource allocation algorithms, there are mainly the following ways to access and uniformly manage communication devices.
[0004] The first is a network access selection algorithm driven by a mathematical model. This network access selection decision-making method uses mathematical model tools to formulate the access network selection problem into an existing mathematical model in the resource allocation problem and derive the optimal algorithm. However, this method has low universality and low feasibility in complex communication scenarios due to the limited and relatively single models used.
[0005] The second is a data-driven network access selection algorithm. This method uses a neural network to learn a large number of communication resource allocation samples, combined with the back-propagation method, to convert the communication resource allocation problem into a matrix multiplication problem. By continuously trial and error correction in a large number of samples, the neural network has the ability to solve fuzzy approximate problems. Such as DNN, CNN, DQN, etc. However, neural networks rely on a large amount of high-precision training sample data. However, data collection is often very difficult in wireless resource management. It takes a long time to collect enough training data, which often cannot meet the requirements of ultra-low latency communication tasks. Therefore, the time overhead and economic cost of this method will be relatively high.
[0006] From the above, it can be seen that in the existing wireless network communication resource allocation algorithm, it is particularly important to solve the access selection and resource management problems of multi-dimensional massive communication devices and communication tasks. In the above two solutions, the timeliness, feasibility and cost of resource allocation cannot meet the actual needs of future space-ground integrated network communications. Summary of the invention
[0007] In order to solve the above problems existing in the prior art, the present invention provides a knowledge-driven resource scheduling method in the Internet of Vehicles in an air-ground integrated scenario. The technical problem to be solved by the present invention is achieved through the following technical solutions:
[0008] The present invention provides a knowledge-driven resource scheduling method in a vehicle network in an air-ground-space integrated scenario, comprising:
[0009] Step S1, obtaining an application scenario of an air-space-ground integrated network, where the application scenario of the air-space-ground integrated network includes multiple vehicles, multiple drones, multiple base stations, and at least one low-Earth orbit satellite;
[0010] Step S2, taking the minimum total delay of service transmission for all vehicles as the optimization target, taking the data transmission rate corresponding to the vehicle service type of the vehicle selecting the access channel and the access channel being occupied by only one vehicle as the constraint conditions, and establishing an optimization model;
[0011] The access channel is one of the sub-channels of the access object, and the access object is one of the drone, base station and satellite in the air-ground integrated network;
[0012] Step S3, solving the optimization model by defining a potential function in combination with an asynchronous dominant actor critic algorithm to obtain a specific access channel;
[0013] Step S4, accessing the air-ground integrated network according to the access channel to complete the resource scheduling process of the vehicle.
[0014] Beneficial effects of the present invention:
[0015] The present invention provides a knowledge-driven resource scheduling method in the Internet of Vehicles under the air-ground-integrated scenario, which uses the knowledge of spatiotemporal correlation to select access channels to solve the interference of adjacent channels, and uses the mathematical knowledge related to communication to define the potential function in combination with the asynchronous dominant actor critic algorithm to solve the optimization model to solve the reward sparsity problem in the early stage of reinforcement learning. At the same time, the subnet learning rate of the asynchronous dominant actor critic algorithm is adjusted, the network convergence speed is accelerated, the reward value learned by the network is improved, and the interpretability and generalization of the neural network are enhanced. The present invention can optimize the resource scheduling and on-demand selection of the vehicle transmission process, and is conducive to the deployment and promotion of vehicle access selection in the air-ground-integrated scenario.
[0016] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 It is a flow chart of a knowledge-driven resource scheduling method in a vehicle network in an air-ground-space integration scenario provided by the present invention;
[0018] Figure 2 It is a model diagram of a vehicle network access system in an air-ground-integrated scenario provided by the present invention;
[0019] Figure 3 It is a graph of different learning rate attenuation results of the subnet provided by the present invention;
[0020] Figure 4 It is a graph of experimental results provided by the present invention that introduces all knowledge;
[0021] Figure 5 It is the relationship between different CPU frequencies and delays provided by the present invention;
[0022] Figure 6 This is the relationship between different transmission powers and delays provided by the present invention. DETAILED DESCRIPTION
[0023] The present invention is further described in detail below with reference to specific embodiments, but the embodiments of the present invention are not limited thereto.
[0024] like Figure 1 As shown, the present invention provides a knowledge-driven resource scheduling method in a vehicle network in an air-ground integrated scenario, including:
[0025] Step S1, obtaining an application scenario of an air-space-ground integrated network, where the application scenario of the air-space-ground integrated network includes multiple vehicles, multiple drones, multiple base stations, and at least one low-Earth orbit satellite;
[0026] refer to Figure 2 , the air-space-ground integrated network includes K vehicles, and n vehicles can optionally access the air-space-ground integrated network, including N drones, M base stations and 1 LEO. By deploying ground base stations (BS), drones (UAVs) and low-orbit satellites (LEOs), full coverage of the air-space-ground integrated network (SAGIN) is achieved to provide stable transmission services for vehicles. Vehicles need to choose to access the network according to business needs to reduce the transmission delay of vehicle-borne services. At the same time, it is assumed that each vehicle will have a new task arriving in each time slot, and the vehicle's computing power is limited, so the vehicle is also required to transmit the task to the cloud for processing. However, the data transmission rate may be lower than the data generation rate, so the vehicle needs to have a storage unit to cache tasks. In the present invention, it is also assumed that each vehicle is equipped with only one antenna, that is, each vehicle can only send one task at a time. When a new task arrives at the vehicle, the vehicle first needs to store the task in the local storage unit. All tasks in the storage unit form a queue, and the first task to arrive will be transmitted first. If the newly generated task queue is too long, even if the task is transmitted at the highest rate, the processing delay will be higher than the delay requirement, and the task will be discarded. Use α k,n,f(t) indicates that the kth vehicle selects the nth access network in time slot t. At this time, assuming that at the beginning of each time slot, the kth vehicle selects the nth access network, and the new task arriving at the vehicle is A k (t), the size of each task is ρ. The present invention defines U k (t) The amount of task data transmitted by the k-th vehicle in time slot t. At the same time, the amount of data stored locally in the k-th vehicle is set to queue Q k (t).
[0027] Step S2, taking the minimum total service transmission delay of all vehicles as the optimization target, taking the data transmission rate corresponding to the vehicle service type of the vehicle selecting the access channel and the access channel being occupied by only one vehicle as the constraint conditions, and establishing an optimization model;
[0028] The access channel is one of the sub-channels of the access object, and the access object is one of the drone, base station and satellite in the air-ground integrated network;
[0029] Since the network access selection for each time slot will seriously affect the service transmission delay, the optimization goal of the algorithm research of the present invention is to minimize the total delay of the vehicle, and the optimization factor is the selection set of the access network. The problem can be expressed as P1:
[0030]
[0031]
[0032]
[0033] τ k (t) is the total service transmission delay when the kth vehicle chooses to access network n in time slot t. k,n,f (t) indicates that the k-th vehicle selects the f-th subchannel of the n-th access network in the t-th time slot. J is defined as the set The number of subchannels. k (t) represents the minimum data transmission rate of the kth vehicle in the tth time slot. m is the task type, type1, type2, and type3 represent delay-sensitive tasks, rate-sensitive tasks, and simple tasks, respectively. P1 is the objective function that minimizes the average time delay of all acquisition tasks of K vehicles within T epochs. Constraint C1 means that each subchannel of the access network can be assigned to at most one vehicle, and constraint C2 limits the minimum data transmission rate of the kth vehicle.
[0034] Step S3, solving the optimization model by defining a potential function in combination with an asynchronous dominant actor critic algorithm to obtain a specific access channel;
[0035] In a specific implementation, step S3 includes:
[0036] S31, determining a pending access object covering the target vehicle in the air-ground integrated network;
[0037] S32, selecting a pending access channel from the sub-channels in the order of optimal and suboptimal according to the idle state of the sub-channel of the pending access object;
[0038] Among them, the optimal state is that its own state is idle and both adjacent subchannels are idle; the suboptimal state is that its own state is idle and one of the adjacent subchannels is idle; if the optimal and suboptimal subchannels are selected, then one of the remaining subchannels is randomly selected as the access channel to be selected.
[0039] The present invention introduces the knowledge of time-space correlation to realize channel selection. Channel selection refers to the adjacent channels of an access network. Due to the signal transmission process, part of the power will be scattered outside the selected channel, that is, adjacent channel leakage occurs, and the value is generally 1% to 5% of the transmission power. When the transmission power is small, the influence of adjacent channel interference can be ignored, but too small a transmission power will affect the network convergence time. Therefore, the present invention must consider the influence of adjacent channel interference on the vehicle network access selection process.
[0040] In order to reduce adjacent channel interference, the present invention uses the knowledge of time-space correlation and proposes a channel selection method, that is, when a vehicle selects a channel, it will select a channel whose adjacent channels are all idle. Until all channels that meet the requirements are selected, a channel is randomly selected from the remaining idle channels where there must be an adjacent channel that produces adjacent channel interference for access.
[0041] S33, respectively establishing pending links between the vehicles and the pending access channels, and calculating the signal-to-noise ratio of each pending link;
[0042] Depending on the pending access object, the pending links can be divided into the following categories:
[0043] (1) Vehicle-UAV Link
[0044] When a vehicle selects a drone network for access, the vehicle-drone data transmission model is used. Because the high-speed mobility of vehicles and drones can cause rapid changes in channel states, the present invention ignores small-scale fading and only considers large-scale channel fading. The path loss of the vehicle-drone link can be calculated and defined as:
[0045]
[0046] Among them, r k (t) represents the horizontal distance between the vehicle and the access network, h k (t) represents the flight altitude of the UAV, v k (t) represents the carrier frequency, c represents the speed of light, and They represent the additional loss in addition to the free space path loss of the line-of-sight (LOS) link and the non-line-of-sight link, respectively, and are variables determined by environmental information. And the probability of LOS in the vehicle-UAV link is:
[0047]
[0048] Among them, a and b are also variables determined by environmental information. Therefore, when the kth vehicle selects the fth subchannel of the nth access network at time slot t, the vehicle-UAV link signal-to-noise ratio is:
[0049]
[0050] in, is the data transmission power of the vehicle, δ 2 is the additive white Gaussian noise power, δ interf It represents the out-of-band power leakage caused by inter-carrier interference. Generally, the leakage value is 1% to 5% of the transmit power. Since each subcarrier of the access network can be assigned to at most one vehicle, for the kth vehicle
[0051]
[0052] (2) Vehicle-Base Station Link
[0053] When the kth vehicle selects the fth subchannel of the nth access network in time slot t, the signal-to-noise ratio of the vehicle-base station link is:
[0054]
[0055] Among them, G k,n,f (t) is the uplink channel gain between the vehicle and the base station link in the tth time slot.
[0056] (3) Vehicle-Satellite Link
[0057] For the kth vehicle on the fth subchannel of the nth access network, at time slot t, the signal-to-noise ratio of the vehicle-satellite link can be calculated as:
[0058]
[0059] G G,n,f (t) is the uplink channel gain between the vehicle and the LEO link in the tth time slot, G rsd,f (t) represents the gain of the LEO receiving antenna, L k,fn,f (t) represents the free space loss between LEO and ground base station, L k,rn,f (t) represents the atmospheric attenuation between LEO and ground base station.
[0060] S34, using the Shannon formula, calculating the maximum data transmission rate that each undetermined link can provide to the vehicle;
[0061] S35, using the maximum data transmission rate, calculating the amount of task data that can be transmitted by each pending link;
[0062] S36, using the task data volume and the maximum data transmission rate, calculate the service transmission delay of each vehicle in each time slot;
[0063] Among them, when the pending link is a communication link established between a vehicle and a satellite, the service transmission delay includes an uplink transmission delay from a ground base station to a satellite and a service transmission delay from a vehicle to a ground base station.
[0064] S37, calculating the vehicle's computing delay according to the task data volume, computing complexity, and the vehicle's CPU frequency;
[0065] S38, calculating the queuing delay of the vehicle according to the amount of local cached data of each vehicle and the average task arrival rate in each time slot;
[0066] S39, obtaining a total delay according to the sum of the service transmission delay, the calculation delay and the queuing delay;
[0067] Assume that at the beginning of each time slot, the kth vehicle selects the nth access network, and the new task arriving at the vehicle is A k (t), the size of each task is ρ. The present invention defines U k (t) The amount of task data transmitted by the k-th vehicle in time slot t. At the same time, the amount of data stored locally in the k-th vehicle is set to queue Q k (t). The delays faced by vehicles when accessing the network can be divided into the following categories:
[0068] (1) Queuing Delay
[0069] When the vehicle is driving, it needs to store the services to be transmitted, and new services are constantly arriving, so there will be many tasks waiting to be processed in the local cache area of the vehicle. According to the little theorem, the queuing delay is equal to the ratio of the average queue waiting amount of task data to the average arrival rate. Therefore, the queuing delay of the local cache area of the vehicle can be calculated as:
[0070]
[0071] in, is the average arrival rate of vehicle cache data.
[0072] (2) Calculation delay
[0073] It is equal to the ratio of the size of the task data packet to the computing power of the vehicle. Define λ as the computational complexity, that is, the number of CPU cycles required to process 1 bit of service data. From this, the computational delay of the vehicle in the tth time slot task can be calculated as:
[0074]
[0075] in, is the CPU frequency used for data calculation by the k-th vehicle in the t-th time slot.
[0076] The uplink transmission delay depends on the size of the task data being transmitted, the uplink transmission rate, and the available transmission resources.
[0077] (3) Uplink transmission delay
[0078] 1) Drones and base stations
[0079] The vehicle transmits the vehicle task to the UAV or base station through the orthogonal sub-channel. Assuming that the channel bandwidth is B0, according to the Shannon formula, the data transmission rate of the vehicle in the tth time slot and the amount of task data that can be transmitted are:
[0080] R k,n,f (t) = α k,n,f (t)B0log(1+γ k,n,t ),
[0081]
[0082] U k (t) = min{Q k (t)+ρA k (t), k (t)τ}
[0083] Among them, τ is the time slot length. Then the transmission delay of the vehicle in the tth time slot is:
[0084]
[0085] 2) Satellite
[0086] The uplink transmission delay of the satellite consists of two parts: the delay of the vehicle uplink transmission to the ground base station and the delay of the ground base station uplink transmission to the satellite. The calculation model of the uplink transmission delay of the vehicle service to the ground base station in the tth time slot is the same as that of the vehicle service uplink transmission to the base station. The bandwidth of the satellite link is B c , then the data transmission rate of the link is calculated as:
[0087]
[0088]
[0089] The bandwidth of the satellite link is B c , the second part of the time delay of the uplink transmission from the ground base station to the satellite can also be calculated by Shannon's formula. In summary, at the tth moment, the total time delay of the vehicle choosing to access the network for service transmission is:
[0090]
[0091] S40, obtaining a reward value function according to the total delay and the maximum data transmission rate;
[0092] The reward function r(t of heterogeneous networks and the total delay τ of vehicle service transmission k (t), data transmission rate R k (t) is related. The formula of the reward value function is as follows:
[0093]
[0094] Since each vehicle transmits different services in the same time slot and each service has different requirements, the vehicle rewards need to be weighted. For delay-sensitive services, w1=-1, w2=0.1; for rate-sensitive services, w1=-0.1, w2=1; for ordinary services, w1=-0.1, w2=0.1.
[0095] S41, using an improved asynchronous dominant actor-critic algorithm to solve the reward value function to obtain a specific access channel.
[0096] In a specific implementation, S41 includes:
[0097] S411, design potential function;
[0098] S412, using the potential function to update the reward value function;
[0099] S413, using the asynchronous advantage actor-critic algorithm to solve the updated reward value function to obtain a specific access channel.
[0100] The reward design and learning method adopted by the present invention is to artificially set the reward function. In the original algorithm, even if the service successfully transmits part of the content in the time slot, the reward it receives is still 0, and only when all tasks are transmitted can a positive reward be obtained. This method of exploring optimization strategies by delaying rewards consumes a lot of resources. In order to solve the problem of sparse rewards, the present invention applies reward shaping in the instantaneous reward generated in each time slot. The present invention defines a potential function Ф(τ k (t)), the agent will receive a reward when it drops from a high potential to a low potential. The specific function is as follows:
[0101] r k ′(t)=r k (t)-Φ(τ k (t))+γΦ(τ k (t+1)),
[0102] where τ k (t) is the delay of the kth vehicle at time slot t, and the value of γ in the present invention is 0.9. Among them, the potential function Φ(τ k (t)) is a linear function, which can be expressed as
[0103]
[0104] where τ min is the minimum delay bound, if the delay is less than it, the reward of the system is zero. max and Φ min It is a man-made parameter.
[0105] The knowledge content cited in the present invention is shown in Table 1.
[0106] Table 1 Overview of the knowledge used
[0107] method Knowledge Advantages How to integrate knowledge question Channel Selection Knowledge of spatiotemporal correlation Reduce the action space Improving learning algorithms Adjacent channel interference Reward Reshaping Mathematical knowledge Accelerated convergence Improving learning algorithms Rewards are sparse
[0108] The content in Table 1 covers the knowledge of spatiotemporal correlation and mathematical knowledge related to communication. The knowledge of spatiotemporal correlation is used to improve the modeling process, and the mathematical knowledge related to communication is combined with the asynchronous advantage actor-critic algorithm to solve the reward sparsity problem directly applied by the algorithm.
[0109] S4131, for the Nth subnet in the asynchronous advantage actor-critic algorithm, set its learning rate to 0.9 N ;
[0110] Since each vehicle transmits different services in the same time slot, and each service has different requirements, the vehicle rewards need to be weighted. For delay-sensitive services, w1=-1, w2=0.1; for rate-sensitive services, w1=-0.1, w2=1; for ordinary services, w1=-0.1, w2=0.1. At the same time, the present invention is designed to make the initial learning rate of the subnet decrease in the order of subnet naming to achieve the setting of decreasing subnet learning rate.
[0111] S4132, using the asynchronous advantage actor-critic algorithm after the setting is completed, solve the updated reward value function to obtain a specific access channel.
[0112] The implementation code of the present invention is as follows:
[0113]
[0114] refer to Figure 3-Figure 6 As shown, Figure 3 This is the result of different learning rate decay of the subnet; the initial learning rate of the network is set to 0.0001. Figure 3 It can be seen that the worker whose learning rate decreases exponentially by 0.9 in the subnetwork converges fastest and obtains the largest reward. Therefore, in the subsequent experiments of the present invention, the learning rate is selected to decrease exponentially by 0.9.
[0115] Figure 4 This is the experimental result diagram of introducing all knowledge; the experiment sets the vehicle to transmit delay-sensitive, rate-sensitive, and common services across time slots. Each access network has 6 sub-channels for vehicles to choose from, and the in-band power is 0.95 of the transmission power. The vehicle access process also uses three improved learning methods: reward reshaping, channel selection, and changing learning rate. The final result is as follows Figure 4 As shown in Figure 2, the learning reward value of the knowledge-driven training algorithm is the highest. The results show that integrated knowledge can effectively reduce the convergence delay by about 4% during the vehicle access process.
[0116] Figure 5 It is the relationship between different CPU frequencies and latency; Figure 5 The relationship between different CPU frequencies and rewards is shown. As can be seen from the figure, the latency decreases as the CPU frequency increases. However, the latency does not improve endlessly with the increase of CPU frequency. In this simulation scenario, when the user's calculation frequency rises to a certain value, the latency curve tends to be stable and no longer improves. This is mainly because the total system latency is composed of three parts: queuing delay, calculation delay, and transmission delay. Increasing the calculation frequency can only reduce the calculation delay, but has no effect on the queuing delay and transmission delay. Therefore, when the calculation frequency is high enough, the other two delays in the total delay dominate, and the delay no longer changes with the increase of CPU frequency.
[0117] Figure 6 It is the relationship between different transmission power and delay. Figure 6 The relationship between different transmission powers and rewards is shown. As can be seen from the figure, the delay decreases as the transmission power increases, that is, increasing the transmission power can significantly reduce the total delay. However, increasing the transmission power does not affect the calculation delay and transmission delay. Therefore, when the transmission power increases to a certain extent, the total delay tends to converge and does not change with the increase of the transmission power.
[0118] The present invention provides a knowledge-driven resource scheduling method in the Internet of Vehicles in the space-ground integration scenario, which uses the knowledge of time-space correlation to solve the interference of adjacent channels and the mathematical knowledge related to communication to solve the reward sparsity problem in the early stage of reinforcement learning. At the same time, it adjusts the subnet learning rate of the asynchronous advantage performer-critic algorithm, accelerates the network convergence speed, improves the reward value learned by the network, and enhances the interpretability and generalization of the neural network. The present invention can optimize the resource scheduling and on-demand selection of the vehicle transmission process, and is conducive to the deployment and promotion of vehicle access selection in the space-ground integration scenario.
[0119] In addition, the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Therefore, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the present invention, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0120] Although the present application is described herein in conjunction with various embodiments, in the process of implementing the claimed application, those skilled in the art may understand and implement other variations of the disclosed embodiments by reviewing the drawings, the disclosure, and the appended claims. In the claims, the word "comprising" does not exclude other components or steps, and "a" or "an" does not exclude a plurality of components or steps.
[0121] The above contents are further detailed descriptions of the present invention in combination with specific preferred embodiments, and it cannot be determined that the specific implementation of the present invention is limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, several simple deductions or substitutions can be made without departing from the concept of the present invention, which should be regarded as falling within the protection scope of the present invention.
Claims
1. A knowledge-driven resource scheduling method in a vehicle network under an air-ground integrated scenario, characterized in that: include: Step S1, obtaining an application scenario of an air-space-ground integrated network, where the application scenario of the air-space-ground integrated network includes multiple vehicles, multiple drones, multiple base stations, and at least one low-Earth orbit satellite; Step S2, taking the minimum total delay of service transmission for all vehicles as the optimization target, taking the data transmission rate corresponding to the vehicle service type of the vehicle selecting the access channel and the access channel being occupied by only one vehicle as the constraint conditions, and establishing an optimization model; The access channel is one of the sub-channels of the access object, and the access object is one of the drone, base station and satellite in the air-ground integrated network; Step S3, solving the optimization model by defining a potential function in combination with an asynchronous dominant actor critic algorithm to obtain a specific access channel; Step S4, accessing the air-ground integrated network according to the access channel to complete the resource scheduling process of the vehicle; Step S3 includes: S31, determining a pending access object covering the target vehicle in the air-ground integrated network; S32, selecting a pending access channel from the sub-channels in the order of optimal and suboptimal according to the idle state of the sub-channel of the pending access object; Among them, the best is that the state of itself is idle and the adjacent subchannels are all idle; the second best is that the state of itself is idle and one of the adjacent subchannels is idle; if the best and second best subchannels are selected, then one of the remaining subchannels is randomly selected as the access channel to be selected; S33, respectively establishing pending links between the vehicles and the pending access channels, and calculating the signal-to-noise ratio of each pending link; S34, using the Shannon formula, calculating the maximum data transmission rate that each undetermined link can provide to the vehicle; S35, using the maximum data transmission rate, calculating the amount of task data that can be transmitted by each pending link; S36, using the task data volume and the maximum data transmission rate, calculate the service transmission delay of each vehicle in each time slot; S37, calculating the vehicle's computing delay according to the task data volume, computing complexity, and the vehicle's CPU frequency; S38, calculating the queuing delay of the vehicle according to the amount of local cached data of each vehicle and the average task arrival rate in each time slot; S39, obtaining a total delay according to the sum of the service transmission delay, the calculation delay and the queuing delay; S40, obtaining a reward value function according to the total delay and the maximum data transmission rate; S41, using an improved asynchronous dominant actor-critic algorithm to solve the reward value function to obtain a specific access channel.
2. According to the knowledge-driven resource scheduling method in the Internet of Vehicles in the space-ground integration scenario of claim 1, it is characterized in that: The air-ground integrated network in step 1 includes K vehicles, and n vehicles can be optionally connected to the air-ground integrated network, including N drones, M base stations and 1 LEO.
3. According to the knowledge-driven resource scheduling method in the Internet of Vehicles in the space-ground integrated scenario of claim 2, the optimization model in step 2 is expressed as: In the formula, τ k (t) is the total service transmission delay when the kth vehicle chooses to access network n in time slot t, α k,n,f (t) indicates that the k-th vehicle selects the f-th subchannel of the n-th access network in the t-th time slot, and J is defined as the set The number of subchannels, R k (t) represents the minimum data transmission rate of the kth vehicle in the tth time slot, m is the task type, type1, type2, and type3 represent delay-sensitive tasks, rate-sensitive tasks, and simple tasks, respectively. P1 is the objective function that minimizes the average time delay of all acquisition tasks of K vehicles within T epochs. Constraint C1 indicates that each subchannel of the access network can be assigned to at most one vehicle, and constraint C2 limits the minimum data transmission rate of the kth vehicle.
4. The knowledge-driven resource scheduling method in the Internet of Vehicles in the space-ground integration scenario according to claim 1 is characterized in that: S41 includes: S411, design potential function; S412, using the potential function to update the reward value function; S413, using the asynchronous advantage actor-critic algorithm to solve the updated reward value function to obtain a specific access channel.
5. The knowledge-driven resource scheduling method in the Internet of Vehicles in the space-ground integration scenario according to claim 1 is characterized in that: In S36, when the link to be determined is a communication link established between the vehicle and the satellite, the service transmission delay includes the uplink transmission delay from the ground base station to the satellite and the service transmission delay from the vehicle to the ground base station.
6. The knowledge-driven resource scheduling method in the Internet of Vehicles in the space-ground integration scenario according to claim 4 is characterized in that: S413 includes: S4131, for the Nth subnet in the asynchronous advantage actor-critic algorithm, set its learning rate to 0.9 N ; S4132, using the asynchronous advantage actor-critic algorithm after the setting is completed, solve the updated reward value function to obtain a specific access channel.
Citation Information
Patent Citations
Network resource allocation method based on space-air-ground integrated system
CN112583566A
Wireless communication network resource allocation algorithm for on-demand dynamic adjustment
CN114219074A