基于多智能体强化学习的车联网多跳卸载和资源分配方法

By employing the MADDPG algorithm based on multi-agent reinforcement learning and a dynamic agent management mechanism, the problem of dynamic optimization of multi-hop offloading and resource allocation in the Internet of Vehicles (IoV) is solved, achieving low-latency task decision-making and efficient resource allocation, and is suitable for IoV edge computing scenarios.

CN122179841BActive Publication Date: 2026-07-17SOUTHEAST UNIV

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SOUTHEAST UNIV
Filing Date
2026-05-12
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In dynamic vehicle-to-everything (V2X) environments, multi-hop offloading and resource allocation face challenges such as rapid changes in network topology due to high-speed vehicle movement, unstable link connections, difficulties in handling mixed action spaces, and the inability of traditional methods to optimize in real time.

Method used

We employ a multi-agent deep deterministic policy gradient (MADDPG) algorithm, combined with Gumbel-Softmax and Sigmoid functions to handle the mixed action space, and design a dynamic agent management mechanism. We adopt a centralized training and distributed execution paradigm to achieve multi-vehicle collaborative decision-making.

Benefits of technology

It achieves joint optimization of multi-hop offloading and resource allocation in dynamic vehicle networking scenarios, reduces the average service latency of tasks, improves the task completion rate, adapts to network dynamics, and reduces computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122179841B_ABST
    Figure CN122179841B_ABST
Patent Text Reader

Abstract

本发明公开了基于多智能体强化学习的车联网多跳卸载和资源分配方法,包括:S1:任务车辆观测本地状态,得到候选服务车辆集合;S2:构建基于深度强化学习的多智能体协同决策框架,任务车辆作为智能体。Actor网络根据本地状态,产生包括服务车辆选择、信道分配和功率控制的混合动作,Critic网络根据全局状态与全局动作,产生期望奖励值;S3:引入Gumbel‑Softmax对离散动作连续化,采用Sigmoid约束连续动作;S4:设计动态智能体管理机制,通过智能体池和掩码向量适应车辆动态变化,实现参数共享,保证训练稳定性;S5:训练完成后,任务车辆根据本地状态,通过Actor网络产生分布式动作并执行。
Need to check novelty before this filing date? Find Prior Art