Vehicle-mounted V2V direct communication channel resource allocation method based on edge calculation

By combining edge computing and deep reinforcement learning, a two-layer Dueling DQN decision model is constructed, which solves the problems of high latency, large interference and insufficient reliability in vehicle-to-vehicle (V2V) communication, realizes efficient resource allocation and seamless switching, and improves the overall performance of the V2V system.

CN121126451AInactive Publication Date: 2025-12-12TIANSHUI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511428625.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-02
Publication Date
2025-12-12
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing channel resource allocation schemes for vehicle-to-everything (V2V) communication suffer from high latency, significant interference, and insufficient reliability, making it difficult to meet the service requirements of low latency and high reliability. Furthermore, they fail to effectively coordinate and optimize resources and suppress interference coupling between V2V and V2I links.

Method used

By employing a two-layer architecture based on edge computing combined with deep reinforcement learning, a Dueling DQN decision model is constructed. Through joint optimization of channel and power, resource allocation is dynamically adjusted in real time to achieve interference suppression and seamless switching of V2V and V2I links, ensuring the QoS of high-priority services.

Benefits of technology

It significantly improves the performance of the vehicle-mounted V2V direct communication system, ensures the high efficiency and stability of safety-related services, adapts to network topology changes brought about by vehicle mobility, avoids communication interruptions, and improves the adaptability and reliability of resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121126451A_ABST
    Figure CN121126451A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of vehicle-mounted communication, and discloses a vehicle-mounted V2V direct communication channel resource allocation method based on edge computing, which comprises the following steps: constructing a double-layer edge communication architecture of an RSU and a vehicle-mounted edge gateway, and defining a V2I / V2V link set; collecting multi-dimensional data and preprocessing the multi-dimensional data; defining an action space and a state space, constructing a Duelling DQN decision model, and initializing a Q network, a target Q network and an experience pool; inputting multi-dimensional data into the trained Q network to obtain an optimal sub-channel and power, verifying SINR and capacity constraints, and issuing an instruction; monitoring an issued resource state, and triggering power / channel adjustment according to real-time interference / time delay data; and evaluating the resource allocation effect and the switching effect to re-execute the model training. According to the invention, channel and power joint optimization is realized through combination of a double-layer edge architecture and deep reinforcement learning, so that mutual interference between V2V and V2I links is suppressed, high-priority service QoS is guaranteed, and stable and efficient resource support is provided for a vehicle-mounted V2V core function.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of vehicle-mounted communication, in particular to a vehicle-mounted V2V direct connection communication channel resource allocation method based on edge computing. BACKGROUND

[0002] Vehicle-mounted V2V (vehicle-to-vehicle) direct connection communication is a key technology for intelligent networked vehicles to realize core functions such as cooperative driving and safety warning. It supports low-latency and high-reliability services such as collision avoidance and vehicle platoon coordination, and has important significance for improving road traffic safety and traffic efficiency.

[0003] Currently, the channel resource allocation of V2V communication faces many challenges. Traditional centralized resource allocation schemes rely on macro base stations as the decision center and need to collect network-wide channel state information (CSI) to develop allocation strategies. However, there is a transmission delay between base stations and vehicles, and it is difficult to adapt to the dynamic changes in network topology caused by high-speed vehicle movement in real time, resulting in a lag in resource allocation response and failing to meet the low-latency requirements of V2V services. Distributed resource allocation schemes, on the other hand, do not require centralized control but often use autonomous competition mechanisms. Vehicles only select channels and power based on local information, which can easily cause same-frequency interference and reduce resource utilization, and cannot guarantee the quality of service (QoS) of safety-related high-priority services.

[0004] In recent years, edge computing technology has been introduced into the field of vehicle-mounted communication due to its proximity to the terminal and low latency. However, existing V2V resource allocation schemes that incorporate edge computing still have some shortcomings. Some schemes do not fully coordinate channel selection and power control, only optimizing a single type of resource, making it difficult to balance interference suppression and capacity improvement. Some schemes do not consider the interference coupling relationship between V2V links and V2I (vehicle-to-infrastructure) links, resulting in mutual influence between the two types of links. At the same time, there is a lack of resource switching adaptation for vehicles crossing edge node coverage areas, which can easily cause communication interruption or QoS degradation, and cannot meet the continuous and reliable communication needs of vehicle-mounted scenarios.

[0005] In summary, there is an urgent need for a vehicle-mounted V2V direct connection communication channel resource allocation method that can fully leverage the advantages of edge computing, taking into account resource coordination optimization, interference control, and dynamic adaptation capabilities, to address the high latency, high interference, and insufficient reliability issues in existing technologies. SUMMARY

[0006] The present application provides a vehicle-mounted V2V direct connection communication channel resource allocation method based on edge computing, which realizes joint optimization of channels and power through a double-layer edge architecture combined with deep reinforcement learning to suppress mutual interference between V2V and V2I links and guarantee the QoS of high-priority services, and dynamically adjusts and seamlessly switches to adapt to vehicle mobility across edge nodes, providing stable and efficient resource support for vehicle-mounted V2V core functions.

[0007] This invention provides a method for allocating vehicle-to-vehicle (V2V) direct communication channel resources based on edge computing, comprising: Construct a two-layer edge communication architecture between the RSU and the vehicle edge gateway, define the V2I / V2V link set, and configure spectrum resources; According to the dual-layer edge communication architecture, multi-dimensional data of V2V / V2I links are collected at a set period, and the multi-dimensional data is preprocessed; wherein, the multi-dimensional data includes channel status, vehicle dynamic information, and QoS / load information; The action space is defined using the data parameters of the dual-layer edge communication architecture, and the state space is defined using the multi-dimensional data after data preprocessing. An edge-driven Dueling DQN decision model is constructed, and the Q network, target Q network, and experience pool are initialized. The initial Dueling DQN decision model is iteratively trained using training samples generated from the multi-dimensional data, while adjacent edge nodes synchronously fuse Q network parameters to output converged model parameters. The multi-dimensional data collected in real time is input into the trained Q network to obtain the optimal sub-channel and power, and the SINR and capacity constraints are verified. The resource allocation instructions for the verified sub-channel and power are then sent to the vehicle terminal of the V2V link. According to the set period, the resource status is monitored and issued, and power / channel adjustment is triggered based on real-time interference / latency data. When the vehicle crosses the edge coverage, Q network parameters are transmitted to achieve seamless handover. Evaluate the effectiveness of resource allocation and switching. If the probability of V2V latency meeting the set value is less than the first set value or the average capacity of V2I is less than the second set value, adjust the reward function weights of the Dueling DQN decision model and retrain the model.

[0008] Furthermore, the steps of constructing a two-layer edge communication architecture between the RSU and the vehicle edge gateway, defining the V2I / V2V link set, and configuring spectrum resources include: The communication area is divided into several grid cells. In each grid cell, one roadside unit (RSU) is deployed as a fixed edge node, and an on-board edge gateway is installed on the vehicle terminal as a mobile edge node, forming a two-layer edge communication architecture of RSU and on-board edge gateway. Configure the links connecting cellular users (CUEs) and macro base stations (BSs) as a V2I link set. Configure the direct communication links between vehicles as a V2V link set. ; Configure d orthogonal sub-channels and allocate the orthogonal sub-channels to V2I links. The V2V links reuse the uplink spectrum of the V2I links according to the 1-to-1 multiplexing principle. A local database is deployed within each edge node to store link status and resource allocation records. The edge nodes and vehicle terminals use the MQTT protocol for data interaction.

[0009] Further, the step of collecting multi-dimensional data from the V2V / V2I link according to the dual-layer edge communication architecture at a set period, and preprocessing the multi-dimensional data, includes: Channel state information, including the real-time channel gain of the V2V link, is acquired through a reference signal generator built into the edge node. Interference power and the channel gain of the V2I link. ; Vehicle dynamic information and braking status are collected via the vehicle's onboard GPS module and CAN bus. The vehicle dynamic information includes the real-time location of the vehicles at both ends of the V2V link. Real-time speed ; The maximum tolerable latency of the V2V link can be obtained through the application layer interface of the V2V terminal. Remaining transmission time The remaining load of the V2I link is obtained through the load monitoring module of the edge node. Sub-channel selection record of the previous time slot ; Instantaneous channel gain of the V2V link Channel gain of V2I link A sliding window filtering algorithm is used for processing; the real-time speed of the vehicle is analyzed. Exponential smoothing is used for processing; if the interference power of the V2V link at a certain moment... If the value exceeds the preset value, it is determined to be an abnormal value, and the interference power from the previous moment is used. To make a substitution.

[0010] Furthermore, the steps of defining the action space using the data parameters of the dual-layer edge communication architecture, defining the state space using the preprocessed multi-dimensional data, constructing the edge-driven Dueling DQN decision model, and initializing the Q-network, target Q-network, and experience pool include: Each edge node manages a V2V link and is configured as an independent intelligent agent. The decision objective of the agent is to satisfy the V2V link latency constraints. V2I link interference constraints Under the premise of maximizing the total capacity of V2V and V2I links; The state space is defined as a 6-dimensional vector based on the preprocessed multi-dimensional data, and its formula is: ;in, This represents the preprocessed V2V link interference power from the previous time slot. This represents the instantaneous channel gain of the preprocessed V2V link. The total interference of the V2I link after preprocessing. Select a record for the previous time slot sub-channel. For the remaining load of the V2I link, This represents the remaining transmission time of the V2V link. Define Action ,in Selecting a sub-channel The number of orthogonal sub-channels. Define the power level; define the reward function. To balance capacity, latency, and interference; among them, For V2I link capacity, For V2V link capacity, These are the weighting coefficients; The Q-function structure is defined using the centralized dominance function of Dueling DQN, as shown in the formula: ;in, For public network parameters, The output layer parameters are the advantage function. Output layer parameters for value functions; Initialize the experience pool in each edge node And set the capacity for storing samples. Initial parameters of the Q-network and the target Q-network Set to a random value, learning rate It uses the Adam optimizer.

[0011] Further, the steps of inputting the real-time collected multi-dimensional data into the trained Q-network to obtain the optimal sub-channel and power, verifying SINR and capacity constraints, and issuing resource allocation instructions for the verified sub-channels and power to the vehicle terminal of the V2V link include: Get the current preprocessing status of the V2V link The input is then fed into a converged Q-network, which outputs the Q-values ​​for all actions. Choose the action with the highest Q value. ,Right now ; V2I Link SINR Verification: Calculate the signal-to-noise ratio (SINR) of the V2I link using the following formula: ,in, Fixed transmit power for V2I links, For spectrum reuse metrics, the verification conditions are set as follows: ; V2V Link SINR and Capacity Verification: The SINR and capacity of a V2V link are calculated using the following formulas: , Set the verification conditions as follows and ; If V2I link SINR verification and V2V link SINR and capacity verification fail, then the action with the second largest Q value is selected. Re-verify until the constraints are met; Edge nodes connect sub-channels via PC5 direct connection interfaces. ,power Resource allocation instructions are sent to the vehicle terminals at both ends of the V2V link.

[0012] Furthermore, the steps of monitoring the resource status according to a set period, triggering power / channel adjustments based on real-time interference / delay data, and transmitting Q network parameters to achieve seamless handover when a vehicle crosses edge coverage include: Edge nodes monitor the interference power of the V2V link in two time slots. Remaining transmission time and vehicle location , and when At that time, the power level of the V2V link will be reduced by one level, and the channel will be switched to the suboptimal sub-channel. ;when At that time, the power level of the V2V link was increased to 23dBm to prioritize ensuring that the transmission rate meets the latency constraints; after adjustment, the verification steps were re-executed to ensure that the constraints meet the verification conditions. When the real-time location of the V2V link vehicle When the current edge node's coverage area is exceeded, the current edge node packages the historical state sequence of the V2V link. Post-training Q-network parameters Allocated sub-channels The data packet is sent to the target edge node through a cross-node coordination protocol; after receiving the data packet, the target edge node analyzes it based on the historical state sequence. Initialize local state, call Q network parameters Perform resource allocation to seamlessly take over decision-making.

[0013] Furthermore, the step of evaluating the resource allocation and switching effects, if the V2V latency satisfies a probability less than a first preset value or the V2I average capacity is less than a second preset value, adjusting the reward function weights of the Dueling DQN decision model and re-executing model training, includes: Each edge node evaluates the resource allocation and handover effects every 1000 communication time slots and generates a performance evaluation report. The evaluation metrics include the average capacity of V2I links, the probability of V2V links meeting latency constraints, and edge decision latency. like and / or average capacity of V2I links If the performance fails to meet the standard, the latency penalty will be increased. With interference penalty coefficient Retrain the model, and after training, use the new Q-network parameters. Replace the original parameters, and then re-enter real-time allocation.

[0014] The present invention also provides a vehicle-to-everything (V2V) direct communication channel resource allocation device based on edge computing, comprising: Define the module to build a two-layer edge communication architecture between the RSU and the vehicle edge gateway, define the V2I / V2V link set, and configure spectrum resources; The acquisition module is used to acquire multi-dimensional data of the V2V / V2I link according to the dual-layer edge communication architecture at a set period, and to preprocess the multi-dimensional data; wherein, the multi-dimensional data includes channel status, vehicle dynamic information and QoS / load information; The construction module is used to define the action space with the data parameters of the dual-layer edge communication architecture, define the state space with the multi-dimensional data after data preprocessing, construct the edge-driven Dueling DQN decision model, and initialize the Q network, target Q network and experience pool. The training module is used to iteratively train the initialized Dueling DQN decision model based on the training samples generated from the multi-dimensional data, while simultaneously fusing Q network parameters with adjacent edge nodes to output converged model parameters. The allocation module is used to input the multi-dimensional data collected in real time into the trained Q network to obtain the optimal sub-channel and power, verify SINR and capacity constraints, and send the resource allocation instructions of the verified sub-channel and power to the vehicle terminal of the V2V link. The adjustment module is used to monitor the resource status issued according to a set period, trigger power / channel adjustment based on real-time interference / latency data, and transmit Q network parameters to achieve seamless handover when the vehicle crosses the edge coverage. The evaluation module is used to evaluate the effectiveness of resource allocation and switching. If the probability of V2V latency meeting the requirement is less than the first set value or the average capacity of V2I is less than the second set value, the reward function weight of the Dueling DQN decision model is adjusted and the model training is re-executed.

[0015] The present invention also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the above-described method.

[0016] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the above-described method.

[0017] The beneficial effects of this invention are as follows: This invention constructs a fixed + mobile dual-layer edge architecture, fully leveraging the advantages of edge nodes' proximity to terminals and low latency. It combines a deep reinforcement learning model to jointly optimize channel selection and power control, considering the interference coupling between V2V and V2I links. This suppresses mutual interference between the two types of links while ensuring the quality of service requirements of high-priority safety services. Furthermore, a real-time dynamic adjustment mechanism adapts to network topology changes caused by vehicle movement, along with parameter coordination and seamless switching across edge nodes, preventing communication interruptions or sudden drops in service quality when vehicles cross coverage areas. Finally, a closed loop is formed based on performance evaluation and iterative model optimization, continuously improving the adaptability of resource allocation and communication reliability. This provides stable and efficient channel resource support for core functions of intelligent connected vehicles, such as collaborative driving and collision warning, significantly enhancing the overall performance of the in-vehicle V2V direct communication system. Attached Figure Description

[0018] Figure 1 This is a schematic diagram of a method flow according to an embodiment of the present invention.

[0019] Figure 2 This is a schematic diagram of the device structure according to an embodiment of the present invention.

[0020] Figure 3 This is a schematic diagram of the internal structure of a computer device according to an embodiment of the present invention.

[0021] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0022] It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0023] like Figure 1 As shown, this invention provides a method for allocating vehicle-to-everything (V2V) direct communication channel resources based on edge computing, including: S1. Construct a two-layer edge communication architecture between the RSU and the vehicle edge gateway, define the V2I (Vehicle to Infrastructure) / V2V (Vehicle to Vehicle) link set, and configure spectrum resources.

[0024] S101 defines an urban traffic model that divides the communication area into 433m×250m grid cells. In each grid cell, one Roadside Unit (RSU) is deployed as a fixed edge node with a coverage radius of ≥200m. At the same time, vehicle-mounted edge gateways are installed on 10% of the vehicle terminals (prioritizing convoy lead vehicles and emergency rescue vehicles) as mobile edge nodes, forming a two-layer edge communication architecture of RSU + vehicle-mounted edge gateway to ensure that there are no signal blind spots in V2V direct communication.

[0025] S102, Link and Resource Definition (1) Define the link set: Configure the links connecting cellular user equipments (CUEs) and macro base stations (BS) as V2I (vehicle-to-infrastructure) link sets. Configure the direct vehicle-to-vehicle communication links as a V2V link set. .

[0026] (2) Configure spectrum resources: Configure d=12 orthogonal sub-channels, the carrier frequency of the sub-channel is 2GHz, and the single channel bandwidth is B=1.5MHz; prioritize the allocation of orthogonal sub-channels to V2I links, and V2V links adopt the 1-to-1 multiplexing principle to reuse the uplink spectrum of V2I links, that is, one V2V link reuses only the spectrum resources of one V2I link to avoid co-channel interference.

[0027] (3) Hardware parameters are set: the antenna height of the macro base station (BS) is 25m, the antenna gain is 8dBi, the antenna gain of the vehicle terminal is 3dBi, the vehicle speed is configured as 20km / h~40km / h, and the system noise power is set to... .

[0028] S103. Deploy a local database with a capacity of ≥1GB inside each edge node (RSU / vehicle edge gateway) to store data such as link status and resource allocation records; the edge node and the vehicle terminal use the MQTT protocol for data interaction, and the interaction latency is controlled to ≤5ms.

[0029] S2. Collect multi-dimensional data of V2V / V2I links according to the dual-layer edge communication architecture at a set period, and preprocess the multi-dimensional data; wherein, the multi-dimensional data includes channel status, vehicle dynamic information and QoS / load information.

[0030] S201. Each edge node collects multi-dimensional data according to a 10ms sensing cycle: (1) Channel State Information (CSI) Acquisition: The instantaneous channel gain of the V2V link is acquired through the reference signal generator built into the edge node. (Includes large-scale fading and fast fading; the path loss factor for large-scale fading is -4, and fast fading follows an exponential distribution with a mean of 1), interference power. (Including interference from V2I links to V2V links) Interference from other V2V links ), and the channel gain of the V2I link. The acquisition accuracy must meet the following requirements: channel gain measurement error ≤3%, interference power measurement error ≤5%.

[0031] (2) Vehicle dynamic information collection: Real-time location of vehicles at both ends of the V2V link is collected through the vehicle-mounted GPS module (positioning accuracy ≤1m). Real-time speed The vehicle's braking status is collected via the vehicle's CAN bus (marked as a high-priority link during emergency braking).

[0032] (3) QoS / load information collection: Obtain the maximum tolerable latency of the V2V link through the application layer interface of the V2V terminal. Remaining transmission time The remaining load of the V2I link is obtained through the load monitoring module of the edge node. (Unit: kb) Sub-channel selection record of the previous time slot .

[0033] S202, Multi-dimensional Data Preprocessing: (1) Channel data filtering: Gain of the acquired V2V link channel V2I link channel gain A sliding window filtering algorithm is used to remove transient noise caused by fast fading. The filter window length is set to 5 sensing periods, and the filtering formula is as follows: ;

[0034] in, The smoothing channel gain of the nth V2V link after filtering is denoted as t, which is the current data acquisition time slot and corresponds to the real-time sensing period of the edge node (e.g., acquiring channel data once every t time slot); i is the time slot index, with a value range of [t−W+1, t], representing all historical and current acquisition times from time slot t−W+1 to the current time slot t contained within the window. The original channel gain of the nth V2V link in the i-th time slot; To smooth the channel gain of the m-th V2I link after filtering. The original channel gain of the m-th V2I link in the i-th time slot.

[0035] (2) Dynamic data smoothing: real-time vehicle speed Exponential smoothing is used to ensure that the speed fluctuation error is ≤5%. The smoothing formula is as follows: .

[0036] (3) Outlier removal: If the V2V link interference power collected at a certain moment is... If the interference exceeds the normal range, the value is considered an abnormal value, and the interference power from the previous moment is used. To make a substitution.

[0037] S3. Define the action space using the data parameters of the dual-layer edge communication architecture, define the state space using the multi-dimensional data after data preprocessing, construct the edge-driven Dueling DQN decision model, and initialize the Q network, target Q network, and experience pool.

[0038] S301. Configure each edge node's managed V2V link as an independent intelligent agent. The decision objective of the agent is to satisfy the V2V link latency constraints. V2I link interference constraints Under the premise of maximizing the total capacity of V2V and V2I links.

[0039] S302, Model Definition: (1) State space Based on the multi-dimensional data after data preprocessing in step S202, the state space is defined as a 6-dimensional vector, and its formula is:

[0040] in, The V2V link interference power in the previous time slot after preprocessing in step S202. This represents the instantaneous channel gain of the preprocessed V2V link. The total interference of the V2I link after preprocessing. Select a record for the previous time slot sub-channel. For the remaining load of the V2I link, This represents the remaining transmission time for the V2V link.

[0041] (2) Action space Define action ,in For sub-channel selection, corresponding to the 12 orthogonal sub-channels configured in step S1). Power level The motion space dimension is 12×3=36.

[0042] (3) Reward function Define the reward function To balance capacity, latency, and interference; among them, For V2I link capacity, For V2V link capacity, These are the weighting coefficients, which are 0.4, 0.4, 0.1, and 0.1, respectively.

[0043] (4) Q-function structure: The centralized dominance function of Dueling DQN is adopted, and the formula is as follows:

[0044] in, The parameters are common to the network, and the network structure is an input layer → two hidden layers (64 neurons / layer). The output layer parameters are the advantage function. These are the output layer parameters for the value function.

[0045] S303. Initialize the experience pool in each edge node. And set the capacity to 10 5 Used to store samples Initial parameters of the Q-network and the target Q-network Set to a random value, learning rate It uses the Adam optimizer.

[0046] S4. The initial Dueling DQN decision model is iteratively trained using training samples generated from the multi-dimensional data. At the same time, adjacent edge nodes synchronously fuse Q network parameters and output converged model parameters.

[0047] Each edge node independently performs model training, with 5000 training iterations (simulation verified as convergent). Adjacent edge nodes (with overlapping coverage ≥ 50m) synchronize Q-network parameters θ via direct fiber optic links at a period of 200 iterations, with synchronization delay controlled ≤ 5ms. A weighted average algorithm is used for parameter fusion, with the fusion formula being: θ 融合 =0.6θ 本地 +0.4θ 邻节点 This is to improve the stability of resource allocation when vehicles cross edge coverage areas.

[0048] S5. Input the multi-dimensional data collected in real time into the trained Q network to obtain the optimal sub-channel and power, verify SINR and capacity constraints, and send the resource allocation instructions for the verified sub-channel and power to the vehicle terminal of the V2V link.

[0049] In each communication time slot (duration 1ms), the edge node performs the following resource allocation operations: S501, Status Input and Action Decision: Get the current preprocessing status of the V2V link The data is then fed into a converged Q-network; the Q-values ​​for all actions are output through the Q-network. Choose the action with the highest Q value. ,Right now ; S502, Interference and Capacity Verification: (1) V2I Link SINR Verification: The signal-to-noise ratio (SINR) of the V2I link is calculated based on the hardware parameters in step S1 and the preprocessed data in step S2. The formula is as follows:

[0050] in, Fixed transmit power for V2I links, The verification condition is set as follows: (1 for reuse, 0 otherwise) .

[0051] (2) V2V Link SINR and Capacity Verification: Calculate the SINR and capacity of the V2V link using the following formula: ;

[0052] Among them, the verification conditions are set as follows: and .

[0053] If V2I link SINR verification and V2V link SINR and capacity verification fail, then the action with the second largest Q value is selected. Re-verify until the constraints are met.

[0054] S503, edge nodes connect to the PC5 direct connection interface to the sub-channel ,power Resource allocation instructions are sent to the vehicle terminals at both ends of the V2V link, with a transmission delay control of ≤3ms.

[0055] S6. Monitor the resource status issued according to the set period, trigger power / channel adjustment based on real-time interference / delay data, and transmit Q network parameters when the vehicle crosses the edge coverage to achieve seamless handover.

[0056] S601, the edge node monitors the interference power of the V2V link in two time slots (2ms in length). Remaining transmission time and vehicle location And trigger the following adjustments: (1) When At this time, the power level of the V2V link is reduced by one level (e.g., from 23dBm to 17dBm), and the channel is switched to the suboptimal subchannel. .

[0057] (2) When At that time, the power level of the V2V link is increased to 23dBm to prioritize ensuring that the transmission rate meets the latency constraint; after adjustment, the verification step S502 is re-executed to ensure that the constraint meets the verification conditions.

[0058] S602, Real-time location of V2V link vehicles When the distance exceeds the coverage area of ​​the current edge node (≥200m from the edge node), a handover process is triggered (handover latency is controlled to ≤8ms to ensure uninterrupted V2V communication): The current edge node packages the historical state sequence of this V2V link. Post-training Q-network parameters Allocated sub-channels The data packet is sent to the target edge node (the coverage area the vehicle is about to enter) via the cross-node cooperation protocol in step S4; after receiving the data packet, the target edge node analyzes it based on the historical state sequence. Initialize local state, call Q network parameters Perform resource allocation to seamlessly take over decision-making.

[0059] S7. Evaluate the effectiveness of resource allocation and switching. If the probability of V2V latency meeting the set value is less than the first set value or the average capacity of V2I is less than the second set value, adjust the reward function weights of the Dueling DQN decision model and re-execute model training.

[0060] S701. Each edge node evaluates resource allocation and handover effectiveness and generates a performance evaluation report based on a period of 1000 communication time slots (1 second in length). The evaluation metrics include: ① Average V2I link capacity: ;in, ① The individual capacity of the m-th V2I link; ② The probability that the V2V link satisfies the delay constraint: P 时延 = Satisfies T max ≤T0 number of V2V links / N max ③ Edge decision latency: The total latency from state perception to instruction issuance.

[0061] S702, if and / or average capacity of V2I links If the performance is deemed unsatisfactory, the following optimization operations will be performed: ① Adjust the weight of the reward function in step S3: (Delay penalty coefficient) and (Interference penalty coefficient) is increased to 0.2; ② Local training in step S4 is re-executed, with the number of iterations reduced to 1000 to quickly adapt to scene changes; ③ After training, the new Q-network parameters are used. Replace the original parameters and re-enter the real-time allocation process in step S5.

[0062] like Figure 2 As shown, the present invention also provides a vehicle-to-everything (V2V) direct communication channel resource allocation device based on edge computing, comprising: Define module 1 to build a two-layer edge communication architecture between the RSU and the vehicle edge gateway, define the V2I / V2V link set, and configure spectrum resources; The acquisition module 2 is used to acquire multi-dimensional data of the V2V / V2I link according to the dual-layer edge communication architecture at a set period, and to preprocess the multi-dimensional data; wherein, the multi-dimensional data includes channel status, vehicle dynamic information and QoS / load information; Module 3 is used to define the action space with the data parameters of the dual-layer edge communication architecture, define the state space with the multi-dimensional data after data preprocessing, construct the edge-driven Dueling DQN decision model, and initialize the Q network, target Q network and experience pool. Training module 4 is used to iteratively train the initialized Dueling DQN decision model based on the training samples generated from the multi-dimensional data, while simultaneously fusing Q network parameters at adjacent edge nodes and outputting converged model parameters. The allocation module 5 is used to input the multi-dimensional data collected in real time into the trained Q network to obtain the optimal sub-channel and power, verify SINR and capacity constraints, and send the resource allocation instructions of the verified sub-channel and power to the vehicle terminal of the V2V link. The adjustment module 6 is used to monitor the resource status issued according to a set period, trigger power / channel adjustment based on real-time interference / delay data, and transmit Q network parameters to achieve seamless switching when the vehicle crosses the edge coverage. Evaluation module 7 is used to evaluate the effectiveness of resource allocation and switching. If the probability of V2V latency is less than the first set value or the average capacity of V2I is less than the second set value, the reward function weight of the Dueling DQN decision model is adjusted and the model training is re-executed.

[0063] In one embodiment, defining module 1 includes: The segmentation unit is used to divide the communication area into several grid units. In each grid unit, one roadside unit (RSU) is deployed as a fixed edge node, and an on-board edge gateway is installed on the vehicle terminal as a mobile edge node, forming a two-layer edge communication architecture of RSU and on-board edge gateway. The first configuration unit is used to configure the links connecting cellular users (CUEs) and macro base stations (BS) as a V2I link set. Configure the direct communication links between vehicles as a V2V link set. ; The second configuration unit is used to configure d orthogonal sub-channels and allocate the orthogonal sub-channels to the V2I link. The V2V link uses the uplink spectrum of the V2I link to reuse the uplink spectrum of the V2I link according to the 1-to-1 multiplexing principle. The deployment unit is used to deploy a local database within each edge node to store link status and resource allocation records. The edge nodes and vehicle terminals use the MQTT protocol for data interaction.

[0064] In one embodiment, the acquisition module 2 includes: The first acquisition unit is used to acquire channel state information, including the real-time channel gain of the V2V link, through the reference signal generator built into the edge node. Interference power and the channel gain of the V2I link. ; The second acquisition unit is used to acquire vehicle dynamic information and vehicle braking status through the vehicle-mounted GPS module and vehicle CAN bus. The vehicle dynamic information includes the real-time location of the vehicles at both ends of the V2V link. Real-time speed ; The acquisition unit is used to obtain the maximum tolerable latency of the V2V link through the application layer interface of the V2V terminal. Remaining transmission time The remaining load of the V2I link is obtained through the load monitoring module of the edge node. Sub-channel selection record of the previous time slot ; A preprocessing unit is used to perform real-time channel gain analysis on the V2V link. Channel gain of V2I link A sliding window filtering algorithm is used for processing; the real-time speed of the vehicle is analyzed. Exponential smoothing is used for processing; if the interference power of the V2V link at a certain moment... If the value exceeds the preset value, it is determined to be an abnormal value, and the interference power from the previous moment is used. To make a substitution.

[0065] In one embodiment, building module 3 includes: The third configuration unit is used to configure each V2V link managed by an edge node as an independent intelligent agent. The decision objective of the agent is to satisfy the V2V link latency constraints. V2I link interference constraints Under the premise of maximizing the total capacity of V2V and V2I links; The first defining unit is used to define the state space as a 6-dimensional vector based on the multi-dimensional data after data preprocessing, and its formula is: ;in, This represents the preprocessed V2V link interference power from the previous time slot. This represents the instantaneous channel gain of the preprocessed V2V link. The total interference of the V2I link after preprocessing. Select a record for the previous time slot sub-channel. For the remaining load of the V2I link, This represents the remaining transmission time of the V2V link. The second definition unit is used to define actions. ,in Selecting a sub-channel The number of orthogonal sub-channels. Define the power level; define the reward function. To balance capacity, latency, and interference; among them, For V2I link capacity, For V2V link capacity, These are the weighting coefficients; The third defining unit is used to define the Q-function structure using the centralized dominance function of Dueling DQN, as follows: ;in, For public network parameters, The output layer parameters are the advantage function. Output layer parameters for value functions; An initialization unit is used to initialize the experience pool in each edge node. And set the capacity for storing samples. Initial parameters of the Q-network and the target Q-network Set to a random value, learning rate It uses the Adam optimizer.

[0066] In one embodiment, the allocation module 5 includes: The status acquisition module is used to obtain the preprocessing status of the current V2V link. The input is then fed into a converged Q-network, which outputs the Q-values ​​for all actions. Choose the action with the highest Q value. ,Right now ; The first verification module is used to verify the SINR of the V2I link: It calculates the SINR of the V2I link using the following formula: ,in, Fixed transmit power for V2I links, For spectrum reuse metrics, the verification conditions are set as follows: ; The second verification module is used to verify the SINR and capacity of the V2V link: The SINR and capacity of the V2V link are calculated using the following formula: , Set the verification conditions as follows and ; The selection unit is used to select the action with the second largest Q value when V2I link SINR verification and V2V link SINR and capacity verification fail. Re-verify until the constraints are met; The distribution unit is used by edge nodes to transmit sub-channels via the PC5 direct connection interface. ,power Resource allocation instructions are sent to the vehicle terminals at both ends of the V2V link.

[0067] In one embodiment, the adjustment module 6 includes: The monitoring unit is used by edge nodes to monitor the interference power of the V2V link in two time slots. Remaining transmission time and vehicle location , and when At that time, the power level of the V2V link will be reduced by one level, and the channel will be switched to the suboptimal sub-channel. ;when At that time, the power level of the V2V link was increased to 23dBm to prioritize ensuring that the transmission rate meets the latency constraints; after adjustment, the verification steps were re-executed to ensure that the constraints meet the verification conditions. Processing unit, used for real-time location of V2V link vehicles When the current edge node's coverage area is exceeded, the current edge node packages the historical state sequence of the V2V link. Post-training Q-network parameters Allocated sub-channels The data packet is sent to the target edge node through a cross-node coordination protocol; after receiving the data packet, the target edge node analyzes it based on the historical state sequence. Initialize local state, call Q network parameters Perform resource allocation to seamlessly take over decision-making.

[0068] In one embodiment, the evaluation module 7 includes: The generation unit is used to evaluate the resource allocation and switching effects of each edge node according to a period of 1000 communication time slots and generate a performance evaluation report. Its evaluation indicators include the average capacity of V2I links, the probability of V2V links meeting latency constraints, and edge decision latency. Optimization unit, used when and / or average capacity of V2I links If this occurs, the performance is deemed substandard, and the latency penalty factor is increased. With interference penalty coefficient Retrain the model, and after training, use the new Q-network parameters. Replace the original parameters, and then re-enter real-time allocation.

[0069] Each of the above modules and units is used to perform the respective steps in the above-mentioned edge computing-based vehicle V2V direct communication channel resource allocation method. The specific implementation methods are as described in the above-mentioned method embodiments, and will not be repeated here.

[0070] like Figure 3 As shown, the present invention also provides a computer device, which may be a server, and its internal structure may be as follows: Figure 3 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides the environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The database stores all data required for the process of allocating resources for the V2V direct-connect communication channel based on edge computing. The network interface is used for communication with external terminals via a network connection. When the computer program is executed by the processor, it implements the V2V direct-connect communication channel resource allocation method based on edge computing.

[0071] Those skilled in the art will understand that Figure 3 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer equipment on which the present application is applied.

[0072] An embodiment of this application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements any of the above-described edge computing-based vehicle V2V direct communication channel resource allocation methods.

[0073] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in this application and in the embodiments can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0074] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, apparatus, article, or method that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, apparatus, article, or method. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, apparatus, article, or method that includes that element.

[0075] The above description is merely a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A method for allocating vehicle-to-everything (V2V) direct communication channel resources based on edge computing, characterized in that, include: Construct a two-layer edge communication architecture between the RSU and the vehicle edge gateway, define the V2I / V2V link set, and configure spectrum resources; According to the dual-layer edge communication architecture, multi-dimensional data of V2V / V2I links are collected at a set period, and the multi-dimensional data is preprocessed; wherein, the multi-dimensional data includes channel status, vehicle dynamic information, and QoS / load information; The action space is defined using the data parameters of the dual-layer edge communication architecture, and the state space is defined using the multi-dimensional data after data preprocessing. An edge-driven Dueling DQN decision model is constructed, and the Q network, target Q network, and experience pool are initialized. The initial Dueling DQN decision model is iteratively trained using training samples generated from the multi-dimensional data, while adjacent edge nodes synchronously fuse Q network parameters to output converged model parameters. The multi-dimensional data collected in real time is input into the trained Q network to obtain the optimal sub-channel and power, and the SINR and capacity constraints are verified. The resource allocation instructions for the verified sub-channel and power are then sent to the vehicle terminal of the V2V link. According to the set period, the resource status is monitored and issued, and power / channel adjustment is triggered based on real-time interference / latency data. When the vehicle crosses the edge coverage, Q network parameters are transmitted to achieve seamless handover. Evaluate the effectiveness of resource allocation and switching. If the probability of V2V latency meeting the requirement is less than the first set value or the average capacity of V2I is less than the second set value, adjust the reward function weights of the Dueling DQN decision model and retrain the model.

2. The method for allocating vehicle-to-everything (V2V) direct communication channel resources based on edge computing according to claim 1, characterized in that, The steps of constructing a two-layer edge communication architecture between the RSU and the vehicle edge gateway, defining the V2I / V2V link set, and configuring spectrum resources include: The communication area is divided into several grid cells. In each grid cell, one roadside unit (RSU) is deployed as a fixed edge node, and an on-board edge gateway is installed on the vehicle terminal as a mobile edge node, forming a two-layer edge communication architecture of RSU and on-board edge gateway. Configure the links connecting cellular users (CUEs) and macro base stations (BSs) as a V2I link set. Configure the direct communication links between vehicles as a V2V link set. ; Configure d orthogonal sub-channels and allocate the orthogonal sub-channels to V2I links. The V2V links reuse the uplink spectrum of the V2I links according to the 1-to-1 multiplexing principle. A local database is deployed within each edge node to store link status and resource allocation records. The edge nodes and vehicle terminals use the MQTT protocol for data interaction.

3. The method for allocating vehicle-to-everything (V2V) direct communication channel resources based on edge computing according to claim 2, characterized in that, The step of collecting multi-dimensional data of V2V / V2I links according to the dual-layer edge communication architecture at a set period, and preprocessing the multi-dimensional data, includes: Channel state information, including the real-time channel gain of the V2V link, is acquired through a reference signal generator built into the edge node. Interference power and the channel gain of the V2I link. ; Vehicle dynamic information and braking status are collected via the vehicle's onboard GPS module and CAN bus. The vehicle dynamic information includes the real-time location of the vehicles at both ends of the V2V link. Real-time speed ; The maximum tolerable latency of the V2V link can be obtained through the application layer interface of the V2V terminal. Remaining transmission time The remaining load of the V2I link is obtained through the load monitoring module of the edge node. Sub-channel selection record of the previous time slot ; Instantaneous channel gain of the V2V link Channel gain of V2I link A sliding window filtering algorithm is used for processing; the real-time speed of the vehicle is analyzed. Exponential smoothing is used for processing; if the interference power of the V2V link at a certain moment... If the value exceeds the preset value, it is determined to be an abnormal value, and the interference power from the previous moment is used. To make a substitution.

4. The method for allocating vehicle-to-everything (V2V) direct communication channel resources based on edge computing according to claim 3, characterized in that, The steps of defining the action space using the data parameters of the dual-layer edge communication architecture, defining the state space using the preprocessed multi-dimensional data, constructing the edge-driven Dueling DQN decision model, and initializing the Q-network, target Q-network, and experience pool include: Each edge node manages a V2V link and is configured as an independent intelligent agent. The decision objective of the agent is to satisfy the V2V link latency constraints. V2I link interference constraints Under the premise of maximizing the total capacity of V2V and V2I links; The state space is defined as a 6-dimensional vector based on the preprocessed multi-dimensional data, and its formula is: ;in, This represents the preprocessed V2V link interference power from the previous time slot. This represents the instantaneous channel gain of the preprocessed V2V link. The total interference of the V2I link after preprocessing. Select a record for the previous time slot sub-channel. For the remaining load of the V2I link, This represents the remaining transmission time of the V2V link. Define Action ,in Selecting a sub-channel The number of orthogonal sub-channels. Define the power level; define the reward function. To balance capacity, latency, and interference; among them, For V2I link capacity, For V2V link capacity, These are the weighting coefficients; The Q-function structure is defined using the centralized dominance function of Dueling DQN, as shown in the formula: ;in, For public network parameters, The output layer parameters are the advantage function. Output layer parameters for value functions; Initialize the experience pool in each edge node And set the capacity for storing samples. Initial parameters of the Q-network and the target Q-network Set the learning rate to a random value. It uses the Adam optimizer.

5. The method for allocating vehicle-to-everything (V2V) direct communication channel resources based on edge computing according to claim 4, characterized in that, The steps of inputting the multi-dimensional data collected in real time into the trained Q network to obtain the optimal sub-channel and power, verifying SINR and capacity constraints, and issuing resource allocation instructions for the verified sub-channels and power to the vehicle terminal of the V2V link include: Get the current preprocessing status of the V2V link The input is then fed into a converged Q-network, which outputs the Q-values ​​for all actions. Choose the action with the highest Q value. ,Right now ; V2I Link SINR Verification: Calculate the signal-to-noise ratio (SINR) of the V2I link using the following formula: ,in, Fixed transmit power for V2I links, For spectrum reuse metrics, the verification conditions are set as follows: ; V2V Link SINR and Capacity Verification: The SINR and capacity of a V2V link are calculated using the following formulas: , Set the verification conditions as follows and ; If V2I link SINR verification, V2V link SINR and capacity verification fail, then the action with the second largest Q value is selected. Re-verify until the constraints are met; Edge nodes connect sub-channels via PC5 direct connection interfaces. ,power Resource allocation instructions are sent to the vehicle terminals at both ends of the V2V link.

6. The method for allocating vehicle-to-everything (V2V) direct communication channel resources based on edge computing according to claim 5, characterized in that, The steps of monitoring the resource status according to a set period, triggering power / channel adjustments based on real-time interference / delay data, and transmitting Q network parameters to achieve seamless handover when vehicles cross edge coverage include: Edge nodes monitor the interference power of the V2V link in two time slots. Remaining transmission time and vehicle location , and when At that time, the power level of the V2V link will be reduced by one level, and the channel will be switched to the suboptimal sub-channel. ;when At that time, the power level of the V2V link was increased to 23dBm to prioritize ensuring that the transmission rate meets the latency constraints; after adjustment, the verification steps were re-executed to ensure that the constraints meet the verification conditions. When the real-time location of the V2V link vehicle When the current edge node's coverage area is exceeded, the current edge node packages the historical state sequence of the V2V link. Q-network parameters after training Allocated sub-channels The data packet is sent to the target edge node through a cross-node coordination protocol; after receiving the data packet, the target edge node analyzes it based on the historical state sequence. Initialize local state, call Q network parameters Perform resource allocation to seamlessly take over decision-making.

7. The method for allocating vehicle-to-everything (V2V) direct communication channel resources based on edge computing according to claim 6, characterized in that, The evaluation of resource allocation and switching effects, including adjusting the reward function weights of the Dueling DQN decision model and retraining the model if the V2V latency satisfaction probability is less than a first set value or the V2I average capacity is less than a second set value, includes: Each edge node evaluates the resource allocation and handover effects every 1000 communication time slots and generates a performance evaluation report. The evaluation metrics include the average capacity of V2I links, the probability of V2V links meeting latency constraints, and edge decision latency. like and / or average capacity of V2I links If the performance fails to meet the standard, the latency penalty will be increased. With interference penalty coefficient Retrain the model, and after training, use the new Q-network parameters. Replace the original parameters, and then re-enter real-time allocation.

8. A vehicle-mounted V2V direct communication channel resource allocation device based on edge computing, characterized in that, include: Define the module to build a two-layer edge communication architecture between the RSU and the vehicle edge gateway, define the V2I / V2V link set, and configure spectrum resources; The acquisition module is used to acquire multi-dimensional data of the V2V / V2I link according to the dual-layer edge communication architecture at a set period, and to preprocess the multi-dimensional data; wherein, the multi-dimensional data includes channel status, vehicle dynamic information and QoS / load information; The construction module is used to define the action space with the data parameters of the dual-layer edge communication architecture, define the state space with the multi-dimensional data after data preprocessing, construct the edge-driven Dueling DQN decision model, and initialize the Q network, target Q network and experience pool. The training module is used to iteratively train the initialized Dueling DQN decision model based on the training samples generated from the multi-dimensional data, while simultaneously fusing Q network parameters at adjacent edge nodes and outputting converged model parameters. The allocation module is used to input the multi-dimensional data collected in real time into the trained Q network to obtain the optimal sub-channel and power, verify SINR and capacity constraints, and send the resource allocation instructions of the verified sub-channel and power to the vehicle terminal of the V2V link. The adjustment module is used to monitor the resource status issued according to a set period, trigger power / channel adjustment based on real-time interference / latency data, and transmit Q network parameters to achieve seamless handover when the vehicle crosses the edge coverage. The evaluation module is used to evaluate the effectiveness of resource allocation and switching. If the probability of V2V latency meeting the requirement is less than the first set value or the average capacity of V2I is less than the second set value, the reward function weight of the Dueling DQN decision model is adjusted and the model training is re-executed.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.