基于深度对比强化学习的无线路由优化方法及网络系统

By combining a deep contrastive reinforcement learning model with multi-dimensional routing metrics, the problem of uneven energy consumption in traditional routing algorithms and high resource consumption in deep reinforcement learning algorithms is solved, achieving efficient and adaptive routing optimization on IoT devices.

CN117749692BActive Publication Date: 2026-07-17CHONGQING KELANDA TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHONGQING KELANDA TECH CO LTD
Filing Date
2023-12-26
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Traditional routing algorithms in wireless multi-hop networks suffer from problems such as unreasonable cluster head selection, uneven energy consumption, and uneven path load. Furthermore, routing optimization algorithms based on deep reinforcement learning consume a lot of computational and storage resources on resource-constrained devices and are difficult to adapt to dynamic environments.

Method used

A deep contrastive reinforcement learning model is adopted to improve the model's generalization ability and adaptability through contrastive learning. It combines multi-dimensional routing metrics such as the energy of candidate forwarding nodes, hop count, and buffer queue count, and adopts a centralized training and distributed interaction architecture to reduce the computational complexity of terminal nodes.

Benefits of technology

It achieves efficient and adaptive routing optimization on resource-constrained nodes, improving network lifetime and data transmission reliability, and reducing data acquisition costs and computational complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117749692B_ABST
    Figure CN117749692B_ABST
Patent Text Reader

Abstract

本发明公开了基于深度对比强化学习的无线路由优化方法及网络系统,该方法应用于物联网无线多跳网络和服务器中,服务器上部署有深度对比强化学习模型,网络包括一个汇聚节点和多个无线终端节点,每个终端节点上部署有Actor网络作为分布式的路由决策模型;包括:基于超帧周期,节点入网时从服务器上获取当前最新路由决策模型;在控制周期,该节点基于最新路由决策模型和局部状态向量,生成最优转发节点;在数据传输周期,该节点传输数据给最优转发节点;该节点将在每个超帧周期内采集的经验信息上传至服务器;服务器将经验信息存储至经验池中,及从中抽取部分经验信息并训练深度对比强化学习模型。本发明降低计算量,提高路由选择效果。
Need to check novelty before this filing date? Find Prior Art