A method and system for vehicle networking resource allocation

By using attention mechanism and deep reinforcement learning resource allocation model in the Internet of Vehicles, the problem of channel state changes caused by high-speed vehicle movement is solved, the accurate prediction of channel state information and real-time optimization of resource allocation is achieved, and the reliability and frequency band utilization of the system are improved.

CN115715021BActive Publication Date: 2025-07-11NANJING UNIV OF POSTS & TELECOMM
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211189956.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-28
Publication Date
2025-07-11
Estimated Expiration
2042-09-28

AI Technical Summary

Technical Problem

The existing Internet of Vehicle resource allocation scheme cannot effectively improve band utilization due to the rapid channel state changes and channel estimation errors caused by high-speed vehicle movement, resulting in reduced reliability, reduced system capacity and increased delay, especially in delay-sensitive applications.

Method used

The time series prediction model based on attention mechanism and deep reinforcement learning algorithm are adopted to periodically inform vehicle position and channel state information through V2V and V2I links. Combined with the resource allocation model of deep reinforcement learning, the channel state is predicted and the resource allocation strategy is optimized to ensure the accuracy of channel state information and the real-time resource allocation.

Benefits of technology

It improves the accuracy and real-time resource allocation in the Internet of Vehicles, reduces channel estimation errors, improves the system reliability and frequency band utilization, and meets the needs of delay-sensitive applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115715021B_ABST
    Figure CN115715021B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and system for vehicle network resource allocation. In the vehicle network communication scenario, vehicle environment information is obtained by using a communication link; through the V2V link and the V2I link, the position information, channel state information, and sub-channel selection strategy of the target vehicle are periodically notified to other vehicle terminals or base stations in the vehicle network communication scenario; if the acquisition of the channel state information of the target vehicle at the current moment fails or is incorrect, a prediction model based on the attention mechanism pre-constructed is used to predict the channel state information at the current moment; the remaining load size, remaining delay length of the previous vehicle V2V link, and the predicted channel state information at the current moment are input into a pre-trained resource allocation model to determine the sub-channel selection strategy, and resource allocation is performed according to the sub-channel selection strategy. Advantages: It ensures that relatively accurate channel state information can be obtained, and solves the disadvantage of the high complexity of solving the convex optimization problem in the traditional resource allocation method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a vehicle networking resource allocation method and system, belonging to the technical field of the Internet of Things. Background Art

[0002] Vehicle-to-everything (V2X) technology is an important field of the Internet of Things (IoT). It is one of the important components of Intelligent Transportation Systems (ITS). As one of the key parts of new infrastructure, vehicle networking has gradually become a new popular field of wireless communication. Vehicle networking plays an important role in both intelligent vehicle technologies, such as autonomous driving and intelligent collision avoidance technology, and in-vehicle information devices, such as in-vehicle audio and video systems and in-vehicle road information systems.

[0003] However, the gradually increasing communication rate requirements of terminals in vehicle networking have brought huge challenges to the capacity and resource allocation of vehicle networking. Moreover, in some latency-sensitive application scenarios, such as autonomous driving and vehicle collision avoidance, higher requirements for the stability and reliability of resource allocation in vehicle networking are put forward. Existing solutions mostly use traditional optimization schemes, and after obtaining channel state information, convex optimization and other tools are used to solve the optimal channel allocation matrix. The advantages of this solution are mature technology, clear principle and excellent theoretical effect. But in the actual situation of vehicle networking, vehicle terminals are often in high-speed motion, which will bring problems of rapidly changing channel states and inaccurate channel estimation.

[0004] In vehicle networking scenarios, in order to improve the frequency band utilization rate, V2V links and V2I links are usually multiplexed. However, multiplexing will lead to problems such as reduced reliability, decreased system capacity and increased latency, which are fatal in applications such as V2V collision warning. At the same time, existing solutions use centralized training, that is, data aggregation and training are carried out on the base station side, and the results are sent to each terminal for execution. This training strategy results in higher actual latency and consumes a certain amount of channel resources. Moreover, existing solutions do not consider the problem of channel estimation error caused by the relatively fast change of channel states in vehicle networking, which leads to a decline in training effect. Therefore, it is necessary to find a technical solution that can improve the problems brought by high-speed driving of vehicle terminals. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to overcome the defects of the prior art and provide a vehicle networking resource allocation method and system.

[0006] To solve the above technical problem, the present invention provides a vehicle networking resource allocation method, including:

[0007] In a pre - constructed vehicle - to - everything (V2X) communication scenario, vehicle environment information is obtained using communication links; the V2X communication scenario includes vehicle - to - vehicle (V2V) communication and vehicle - to - infrastructure (V2I) communication, and the communication links include V2V links and V2I links. The V2V links multiplex the uplink of the V2I links, and the uplink of each V2I link is multiplexed by only one V2V link; the vehicle environment information includes the channel state information of each sub - channel, the remaining load size of the current vehicle's V2V link, and the remaining delay length; the channel state information includes channel gain and interference intensity; the sub - channels are the sub - channels divided after the frequency range assigned to the V2I link, and the V2V links multiplex these sub - channels.

[0008] Through the V2V links and V2I links, the position information, channel state information, and sub - channel selection strategy of the target vehicle are periodically reported to other vehicle terminals or base stations in the V2X communication scenario.

[0009] If the acquisition of the channel state information of the target vehicle at the current moment fails or is incorrect, then the channel gain and interference intensity sequences of the target vehicle at n moments are acquired and input into a pre - constructed prediction model based on the attention mechanism to predict the channel state information at the current moment.

[0010] The remaining load size, remaining delay length of the previous vehicle's V2V link, and the predicted channel state information at the current moment are input into a pre - trained resource allocation model to obtain the current resource allocation action. According to the current resource allocation action, a sub - channel selection strategy is obtained, and resource allocation is performed according to the sub - channel selection strategy.

[0011] Furthermore, the prediction model based on the attention mechanism adopts the Seq2Seq algorithm with the attention mechanism.

[0012] Furthermore, the training process of the resource allocation model includes:

[0013] Define the channel gain of the m - th V2I link to the base station on the f - th sub - channel as g m,B [f], the channel gain of the k - th pair of V2V users on the f - th sub - channel as g k [f], the interference gain of the k - th pair of V2V users on the f - th sub - channel by the k'- th pair of V2V users as g k′,k [f], the interference gain of the k - th pair of V2V users on the f - th sub - channel by the m - th V2I user as g m,k [f], the interference gain of the k - th pair of V2V users on the f - th sub - channel to the base station as g k,B [f];

[0014] Define the action space A as a k ={Pk , c k,f} where P k is the transmit power of the k-th pair of V2V users on the subchannel, and c k,f represents the weight of the f-th subchannel for the k-th pair of users;

[0015] Define the state space S as the vehicle environment information of the target vehicle, denoted as s t = {G m,B , G k , I k′,k , I m,k , G k,B , L t , U t}, where G m,B is the channel gain vector of the m-th V2I link to the base station on each subchannel, G k is the channel gain vector of the k-th pair of V2V users on each subchannel, I k′,k is the interference gain vector of the k-th pair of V2V users on each subchannel from the k'-th pair of V2V users, I m,k is the interference gain vector of the k-th pair of V2V users on each subchannel from the m-th V2I user, G k,B is the interference intensity vector of the k-th pair of V2V users on each subchannel to the base station, L t is the remaining allowable delay of the buffer content, and U t is the remaining buffer size;

[0016] Define the reward function R. The goal is to maximize the capacity of the V2I link, ensure the reliability of the V2V link, and meet the delay condition. The reward function R is defined as r t = λ c ∑ m∈M C m,t + λ d ∑ k∈K (γ k,t - γ0) + λ t ∑ k∈K L k,t U k,t , where C m,t is the capacity of the m-th V2I, γ k,t is the SINR of the k-th V2V link, γ0 is the outage SINR threshold, and λ c , λ d , λ t are the weights of the capacity, reliability, and delay condition of the V2I link, respectively;

[0017] During training, the DDPG algorithm is used. The vehicle inputs the state s t at the current moment and obtains the action a k, and store this set of state transition processes in the experience pool; a k It includes a vector and a power value. The vector represents the weight of each sub-channel. The target vehicle sorts according to the weight sizes of the actions a of different vehicles obtained k as its own sub-channel allocation strategy, synchronizes with the V2V pair, selects the sub-channel where the weights of a pair of V2V clients are both in the front and not much different as the finally used sub-channel, and according to a k 's power value generates an action through the Actor network; after executing the action, score according to the reward function R and store it in the experience pool. Based on the scoring result and the experience pool, use soft update and experience replay mechanism to update the network weights. When the training times are reached, update the parameters of the target network to obtain a trained resource allocation model.

[0018] A vehicle networking resource allocation system, including:

[0019] An acquisition module, used to obtain vehicle environment information through a communication link in a pre-constructed vehicle networking communication scenario; the vehicle networking communication scenario is vehicle-to-vehicle communication V2V and vehicle-to-infrastructure communication V2I, and the communication link includes a V2V link and a V2I link. The V2V link multiplexes the uplink of the V2I link, and the uplink of each V2I link is only multiplexed by one V2V link; the vehicle environment information includes the channel state information of each sub-channel, the remaining load size and the remaining delay length of the current vehicle's V2V link; the channel state information includes channel gain and interference intensity; the sub-channel is the sub-channel divided after the frequency range allocated to the V2I link, and the V2V link multiplexes these sub-channels;

[0020] A notification module, used to periodically notify the position information, channel state information, and sub-channel selection strategy of the target vehicle to other vehicle terminals or base stations in the vehicle networking communication scenario through the V2V link and the V2I link;

[0021] A prediction module, used to, if the channel state information of the target vehicle at the current moment is obtained unsuccessfully or incorrectly, obtain the channel gain and interference intensity sequences of the target vehicle at n moments, input them into a pre-constructed prediction model based on the attention mechanism, and predict the channel state information at the current moment;

[0022] An allocation module, used to input the remaining load size, remaining delay length of the previous vehicle's V2V link, and the predicted channel state information at the current moment into a pre-trained resource allocation model to obtain the current resource allocation action, obtain the sub-channel selection strategy according to the current resource allocation action, and perform resource allocation according to the sub-channel selection strategy.

[0023] A computer-readable storage medium storing one or more programs, characterized in that the one or more programs include instructions that, when executed by a computing device, cause the computing device to execute any of the methods described above.

[0024] A computing device, comprising,

[0025] One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the methods described above.

[0026] Advantages achieved by the present invention:

[0027] The present invention uses a time series prediction model based on the attention mechanism to predict the channel gain and interference intensity, ensuring that relatively accurate channel state information can be obtained; the deep reinforcement learning algorithm effectively solves the drawback of high complexity in solving convex optimization problems by traditional resource allocation methods, which is beneficial for in-vehicle devices to use; using time series prediction based on deep learning and the attention mechanism solves the pain points of fast-changing channel states in vehicle-to-everything networks and the difficulty of obtaining real-time channel state information, providing good information support for the resource allocation module. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is a schematic flow chart of the method of the present invention;

[0029] Figure 2 is a schematic module diagram of the system of the present invention. DETAILED DESCRIPTION

[0030] The present invention will be further described below with reference to the accompanying drawings. The following embodiments are only used to more clearly illustrate the technical solutions of the present invention and should not be used to limit the protection scope of the present invention.

[0031] As Figure 1 shown, a vehicle-to-everything network resource allocation method includes:

[0032] Construct a vehicle-to-everything network communication scenario, including vehicle-to-vehicle communication (V2V) and vehicle-to-infrastructure communication (V2I). The V2V link multiplexes the V2I uplink, and each V2I uplink is multiplexed by only one V2V link.

[0033] Initialization phase: The vehicle initializes the deep reinforcement learning algorithm, such as network parameters, experience pool, etc. Initialize the environment perception module, obtain link information and perform the first data synchronization.

[0034] Define the action space A as a k ={P k ,ck,f}, define the state space S as s t = {G m,B , G k , I k′,k , I m,k , G k,B , L t , U t}, define the reward function as r t = λ c ∑ m∈M C m,t + λ d ∑ k∈K (γ k,t - γ0) + λ t ∑ k∈K L k,t U k,t , reasonably adjust the weight factor according to the actual situation of the vehicle to adjust the V2I link capacity and the stability of the V2V link.

[0035] The target vehicle collects vehicle environment information. If the acquisition of the current channel state fails or the acquired value has obvious errors, use the time series prediction model for prediction to obtain the estimated value at the current moment. Convert the value at the current moment into the state vector s t Input it into the resource allocation module.

[0036] The resource allocation model accepts the state vector s t , and uses the DDPG algorithm during training. The vehicle inputs the state s at the current moment t , and applies the obtained action a k to the environment, and stores this set of state conversion processes in the experience pool. Use the soft update and experience replay mechanisms to update the network weights.

[0037] After waiting for the specified time slot, return to step 4 for the next resource allocation calculation.

[0038] When the action a k does not change much, after a set period, use the data synchronization module to synchronize the channel state, location information, sub-channel selection strategy, etc. And update it to the online network parameters.

[0039] Correspondingly, the present invention further provides an acquisition module, which is used to obtain vehicle environment information through a communication link in a pre-constructed vehicle networking communication scenario; the vehicle networking communication scenario is vehicle-to-vehicle communication V2V and vehicle-to-infrastructure communication V2I, and the communication link includes a V2V link and a V2I link. The V2V link multiplexes the uplink of the V2I link, and the uplink of each V2I link is only multiplexed by one V2V link; the vehicle environment information includes the channel state information of each sub-channel, the remaining load size of the current vehicle's V2V link, and the remaining delay length; the channel state information includes channel gain and interference intensity; the sub-channel is a sub-channel divided after the frequency range allocated to the V2I link, and the V2V link multiplexes these sub-channels;

[0040] A notification module, which is used to periodically notify the position information, channel state information, and sub-channel selection strategy of the target vehicle to other vehicle terminals or base stations in the vehicle networking communication scenario through the V2V link and the V2I link;

[0041] A prediction module, which is used to, if the acquisition of the channel state information of the target vehicle at the current moment fails or is incorrect, obtain the channel gain and interference intensity sequences of the target vehicle at n moments, input them into a pre-constructed prediction model based on the attention mechanism, and predict the channel state information at the current moment;

[0042] An allocation module, which is used to input the remaining load size, remaining delay length of the previous vehicle's V2V link, and the predicted channel state information at the current moment into a pre-trained resource allocation model to obtain the current resource allocation action, obtain the sub-channel selection strategy according to the current resource allocation action, and perform resource allocation according to the sub-channel selection strategy.

[0043] As Figure 2 shown, the system specifically includes: an environment perception module, a time series prediction module, a data synchronization module, and a resource allocation module.

[0044] The environment perception module: is used to collect the environment information of the vehicle, including the channel gain of each sub-channel, interference intensity, the remaining load size of the current vehicle's V2V link, and the remaining delay length (specified maximum delay - used delay), etc. This module includes the following processes:

[0045] 1) Environment information acquisition: The vehicle obtains the channel gain of the current vehicle on each sub-channel through a wireless module, which can be blind estimation, semi-blind estimation, and the method of using pilots, etc. At the same time, the vehicle senses the intensity of the interference signal on each channel.

[0046] 2) Position information acquisition: The vehicle obtains the current real-time position through a positioning system.

[0047] 3) Link state acquisition: Collect the current data volume to be sent and the corresponding remaining delay, denoted as L t ,U t .

[0048] 4) Synchronization information acquisition: Collect the channel state information of other terminals from the data synchronization module.

[0049] 5) Data preprocessing: Verify the accuracy of the collected environmental information and link information, and delete the data with obvious errors. If the acquisition of channel information fails, prepare to use the time series prediction module for prediction. Then convert the collected data into vector format for convenient reading by the deep learning algorithm in the resource allocation module.

[0050] Sequence processing module: To address the problem that the channel state information in the vehicle-to-everything (V2X) network changes rapidly and is difficult to obtain in real time, the channel gain sequence is processed using deep learning to predict the channel gain and interference intensity at the current moment. It mainly includes the following steps:

[0051] 1) Collect the channel gain and interference intensity sequences at the previous n moments from the environmental perception module.

[0052] 2) Since there are errors such as noise in the data obtained by the environmental perception module, operations such as weighted average and filtering are performed on the sequence to obtain the optimal sequence value. And extract features such as the periodicity of the sequence for the next prediction.

[0053] 3) When the current measurement value cannot be accurately obtained, use the Seq2Seq algorithm with attention mechanism, such as Transformer, etc. to predict the sequence and obtain the predicted value at the current moment. The weighted average of this value or the predicted value and the measurement value is used as the channel gain or interference intensity at the current moment.

[0054] Data synchronization module: Through the use of vehicle-to-vehicle (V2V) links and vehicle-to-infrastructure (V2I) links, periodically notify other terminals in the V2X network of its own location information, channel state, sub-channel selection strategy, etc. This module mainly includes two synchronization mechanisms: channel state information sharing based on V2V links and synchronization of terminal location information and sub-channel selection strategy between V2X networks based on V2I.

[0055] Channel state information sharing based on V2V links is synchronized between the established V2V terminals. After the power allocation reaches a given period, the respective location information and channel state are notified through the V2V link, and the replacement strategy is also synchronized when the sub-channel needs to be changed. The period here can be set by giving a value, or it can be required to synchronize after each execution of environmental information collection. It can also be performed when the difference between the action vector output by the resource allocation module and the previous vector is greater than the preset threshold, that is, when the allocation strategy changes significantly.

[0056] The synchronization of the terminal location information and sub-channel selection strategy in the vehicle-to-infrastructure (V2I) based Internet of Vehicles (IoV) occurs between the base station and the IoV terminals. The IoV terminals periodically report their location information, channel status, sub-channel selection strategy, etc. to the base station, and the base station forwards them to other terminals. In this process, the resource allocation module of the sub-channel selection strategy gives the priority weight of each sub-channel, and the vehicle makes a choice based on the weight and the selection of the vehicle-to-vehicle (V2V) link between vehicles close to itself, preferentially selecting channels with less interference and no reuse with nearby vehicles. The result is sent to the base station as a sub-channel selection strategy vector.

[0057] By using the data synchronization module, each vehicle can synchronously obtain the channel information of other terminals while acquiring the channel status information around itself, so as to optimize the strategy of V2V link sub-channel selection and share the strategy with the other terminal of the V2V pair.

[0058] Resource allocation module: Using the environmental information and channel status information obtained by the environmental perception module and the time series prediction module, it processes and analyzes the data through deep reinforcement learning to obtain the resource allocation method that maximizes the reward function for the current vehicle. Considering that in the actual IoV application scenario, both V2V and V2I links need to optimize power allocation, so the present invention will jointly optimize the effects of V2V and V2I links, improve the V2I link capacity while ensuring the reliability of the V2V link, and can dynamically adjust the weights of various aspects by adjusting the reward function parameters of the neural network. This module mainly includes the following steps:

[0059] 1) Construct an IoV communication scenario, including vehicle-to-vehicle communication (V2V) and vehicle-to-infrastructure communication (V2I). The V2V link multiplexes the V2I uplink, and each V2I uplink is only multiplexed by one V2V link.

[0060] 2) The sets of V2I links and V2V links are represented by M = {1, 2, …, M} and K = {1, 2, …, K} respectively. In the time domain, the time is divided into time slots with a duration of t, and in the frequency domain, the total system bandwidth W total is divided into F = W total / W sub-channels according to the sub-channel bandwidth W, and each sub-channel becomes a resource block (RB) after one time slot.

[0061] 3) During the training phase, the DDPG algorithm is used with the goal of minimizing the V2V user outage probability and maximizing the V2I user system capacity, while meeting the latency conditions. During each training session, the vehicle sends the locally collected information such as channel gain, interference intensity, remaining load size, and remaining latency information into the Actor network and the Critic network. The Actor network selects actions based on the input, including transmit power and subchannel multiplexing. The Critic network calculates the reward value and updates the network parameters.

[0062] 4) After each round of training is completed, the current V2V user transmit power and channel multiplexing strategy are obtained, including the weight of each subchannel. The vehicle selects subchannels based on the weight of the subchannels and the multiplexing strategies of other links obtained from the synchronization module to avoid repeated multiplexing. And the power is allocated using this strategy.

[0063] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0064] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0065] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0066] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby providing instructions for implementing the steps of the process Figure 1 one process or a plurality of processes and / or blocks Figure 1 or steps for implementing the functions specified in a plurality of blocks.

[0067] The foregoing is only a preferred embodiment of the present invention, and it should be pointed out that for those of ordinary skill in the art, several improvements and modifications can be made without departing from the technical principle of the present invention, and these improvements and modifications should also be regarded as the protection scope of the present invention.

Claims

1. A method for allocating vehicle networking resources, characterized in that Including: In a pre-constructed vehicle networking communication scenario, use a communication link to obtain vehicle environment information; The vehicle networking communication scenario is vehicle-to-vehicle communication V2V and vehicle-to-infrastructure communication V2I. The communication link includes a V2V link and a V2I link. The V2V link multiplexes the uplink of the V2I link, and the uplink of each V2I link is multiplexed by only one V2V link; the vehicle environment information includes the channel state information of each sub-channel, the remaining load size of the current vehicle's V2V link, and the remaining delay length; the channel state information includes channel gain and interference intensity; the sub-channel is a sub-channel divided after the frequency range assigned to the V2I link, and the V2V link multiplexes these sub-channels; Through the V2V link and the V2I link, periodically notify other vehicle terminals or base stations in the vehicle networking communication scenario of the position information, channel state information, and sub-channel selection strategy of the target vehicle; If the acquisition of the channel state information of the target vehicle at the current moment fails or is incorrect, obtain the channel gain and interference intensity sequences of the target vehicle at n moments, and input them into a pre-constructed prediction model based on the attention mechanism to predict the channel state information at the current moment; Input the remaining load size, remaining delay length of the previous vehicle's V2V link, and the predicted channel state information at the current moment into a pre-trained resource allocation model to obtain the current resource allocation action, obtain the sub-channel selection strategy according to the current resource allocation action, and perform resource allocation according to the sub-channel selection strategy; The training process of the resource allocation model includes: Define the channel gain of the m-th V2I link to the base station on the f-th subchannel as g m,B [f], and define the channel gain of the k-th pair of V2V users on the f-th subchannel as g k [f]. Define the interference gain of the k-th pair of V2V users on the f-th subchannel caused by the k'-th pair of V2V users as g k′,k [f]. Define the interference gain of the k-th pair of V2V users on the f-th subchannel caused by the m-th V2I user as g m,k [f]. Define the interference gain of the k-th pair of V2V users on the f-th subchannel to the base station as g k,B [f]; Define the action space A as a k = {P k , c k,f}, where P k is the transmit power of the k-th pair of V2V users on the sub-channel, and c k,f represents the weight of the f-th sub-channel for the k-th pair of users; Define the state space S as the vehicle environment information of the target vehicle, denoted as s t ={G m,B , G k , I k′,k , I m,k , G k,B , L t , U t}, where G m,B is the channel gain vector of the m-th V2I link to the base station on each sub-channel, G k is the channel gain vector of the k-th pair of V2V users on each sub-channel, I k′,k is the interference gain vector of the k-th pair of V2V users on each sub-channel caused by the k'-th pair of V2V users, I m,k is the interference gain vector of the k-th pair of V2V users on each sub-channel caused by the m-th V2I user, G k,B is the interference intensity vector of the k-th pair of V2V users on each sub-channel to the base station, L t is the remaining allowable delay of the buffer content, U t is the remaining buffer size; Define the reward function $R$. The goal is to maximize the capacity of the V2I link while ensuring the reliability of the V2V link and meeting the latency condition. The reward function $R$ is defined as $r$ t $=$ $\lambda$ c $\sum$ m∈M $C$ m,t $+$ $\lambda$ d $\sum$ k∈K $(\gamma$ k,t $-$ $\gamma_0) + \lambda$ t $\sum$ k∈K $L$ k,t $U$ k,t , where $C$ m,t is the capacity of the $m$-th V2I, $\gamma$ k,t is the SINR of the $k$-th V2V link, $\gamma_0$ is the outage SINR threshold, and $\lambda$ c , $\lambda$ d , $\lambda$ t are the weights of the capacity, reliability, and latency condition of the V2I link, respectively; The DDPG algorithm is used for training, and the vehicle inputs the current state s t , the obtained action a k , and store this set of state transition processes in the experience pool; a k Contains a vector and a power value. The vector represents the weight of each subchannel. The target vehicle follows the actions a obtained from different vehicles. k The weight order in the is used as its own sub-channel allocation strategy, and it is synchronized with the V2V pair, and the sub-channels with the highest weights and the smallest difference in the weights of a pair of V2V users are selected as the final sub-channels to be used, and the weights are sorted according to a. k The power value generates an action through the Actor network; after the action is executed, it is scored according to the reward function R and stored in the experience pool. Based on the scoring result and the experience pool, the network weights are updated using soft update and experience playback mechanism. When the number of training times is reached, the parameters of the target network are updated to obtain a trained resource allocation model.

2. The vehicle networking resource allocation method according to claim 1, wherein The prediction model based on the attention mechanism adopts the Seq2Seq algorithm of the attention mechanism.

3. A vehicle networking resource allocation system, characterized in that, Including: An acquisition module, used to obtain vehicle environment information using a communication link in a pre-constructed vehicle networking communication scenario; The vehicle networking communication scenario is vehicle-to-vehicle communication V2V and vehicle-to-infrastructure communication V2I. The communication link includes a V2V link and a V2I link. The V2V link multiplexes the uplink of the V2I link, and the uplink of each V2I link is multiplexed by only one V2V link; the vehicle environment information includes the channel state information of each sub-channel, the remaining load size of the current vehicle's V2V link, and the remaining delay length; the channel state information includes channel gain and interference intensity; the sub-channel is a sub-channel divided after the frequency range assigned to the V2I link, and the V2V link multiplexes these sub-channels; A notification module, used to periodically notify other vehicle terminals or base stations in the vehicle networking communication scenario of the position information, channel state information, and sub-channel selection strategy of the target vehicle through the V2V link and the V2I link; A prediction module, used to obtain the channel gain and interference intensity sequences of the target vehicle at n moments and input them into a pre-constructed prediction model based on the attention mechanism to predict the channel state information at the current moment if the acquisition of the channel state information of the target vehicle at the current moment fails or is incorrect; A distribution module is configured to input the remaining load size, remaining delay length of the previous vehicle V2V link, and the predicted channel state information at the current moment into a pre-trained resource allocation model to obtain a current resource allocation action, obtain a sub-channel selection strategy according to the current resource allocation action, and perform resource allocation according to the sub-channel selection strategy; The training process of the resource allocation model includes: Define the channel gain of the m-th V2I link to the base station on the f-th subchannel as g m,B [f], and define the channel gain of the k-th pair of V2V users on the f-th subchannel as g k [f]. Define the interference gain of the k-th pair of V2V users on the f-th subchannel caused by the k'-th pair of V2V users as g k′,k [f]. Define the interference gain of the k-th pair of V2V users on the f-th subchannel caused by the m-th V2I user as g m,k [f]. Define the interference gain of the k-th pair of V2V users on the f-th subchannel to the base station as g k,B [f]; Define the action space A as a k ={P k , c k,f}, where P k is the transmit power of the k-th pair of V2V users on the sub-channel, and c k,f represents the weight of the f-th sub-channel for the k-th pair of users; Define the state space S as the vehicle environment information of the target vehicle, denoted as s t ={G m,B ,G k ,I k′,k ,I m,k ,G k,B ,L t ,U t}, where G m,B is the channel gain vector of the m-th V2I link to the base station on each sub-channel, G k is the channel gain vector of the k-th pair of V2V users on each sub-channel, I k′,k is the interference gain vector of the k-th pair of V2V users on each sub-channel affected by the k'-th pair of V2V users, I m,k is the interference gain vector of the k-th pair of V2V users on each sub-channel affected by the m-th V2I user, G k,B is the interference intensity vector of the k-th pair of V2V users on each sub-channel on the base station, L t is the remaining allowable delay of the buffer content, U t is the remaining buffer size; Define the reward function \(R\). The goal is to maximize the capacity of the V2I link, while ensuring the reliability of the V2V link and meeting the delay condition. The reward function \(R\) is defined as \(r\) t =\(\lambda\) c \(\sum\) m∈M \(C\) m,t +\(\lambda\) d \(\sum\) k∈K (\(\gamma\) k,t -\(\gamma_0\))+\(\lambda\) t \(\sum\) k∈K \(L\) k,t \(U\) k,t , where \(C\) m,t is the capacity of the \(m\)-th V2I, \(\gamma\) k,t is the SINR of the \(k\)-th V2V link, \(\gamma_0\) is the outage SINR threshold, and \(\lambda\) c , \(\lambda\) d , \(\lambda\) t are the weights of the capacity, reliability, and delay condition of the V2I link, respectively; The DDPG algorithm is used for training, and the vehicle inputs the current state s t , the obtained action a k , and store this set of state transition processes in the experience pool; a k Contains a vector and a power value. The vector represents the weight of each subchannel. The target vehicle follows the actions a obtained from different vehicles. k The weight order in the is used as its own sub-channel allocation strategy, and it is synchronized with the V2V pair, and the sub-channels with the highest weights and the smallest difference in the weights of a pair of V2V users are selected as the final sub-channels to be used, and the weights are sorted according to a. k The power value generates an action through the Actor network; after the action is executed, it is scored according to the reward function R and stored in the experience pool. Based on the scoring result and the experience pool, the network weights are updated using soft update and experience playback mechanism. When the number of training times is reached, the parameters of the target network are updated to obtain a trained resource allocation model.

4. A computer-readable storage medium storing one or more programs, characterized in that, The one or more programs include instructions that, when executed by a computing device, cause the computing device to execute any of the methods according to claims 1 to 2.

5. A computing device, characterized in that, Comprising, One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include instructions for executing any of the methods according to claims 1 to 2.

Citation Information

Patent Citations

  • 5G Internet of Vehicles V2V resource allocation method adopting depth deterministic strategy gradient algorithm

    CN112995951A

  • Channel prediction method and device, electronic equipment and storage medium

    CN114422059A