一种基于网络防御智能体的决策方法

By constructing a network system topology graph and generating agents using sigma rules, deploying attacking and defending agents, and utilizing reinforcement learning training and state encoders for action decisions, the computational complexity and instability issues in multi-agent reinforcement learning are resolved, thereby improving the adaptability and learning stability of the agents.

CN120639464BActive Publication Date: 2026-07-17JIANGSU RUINING XINCHUANG TECH CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
JIANGSU RUINING XINCHUANG TECH CO LTD
Filing Date
2025-07-18
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

In complex network environments, the computational complexity of single-agent reinforcement learning increases with the environmental state and action space size, leading to impractical learning. In multi-agent reinforcement learning, agents cannot utilize information from other agents, resulting in an unstable learning process with poor convergence.

Method used

Construct a network system topology map, generate agents based on sigma rules, match network threats based on threat intelligence, deploy attack and defense agents, train agents through reinforcement learning, make action decisions using agent state encoders and evaluation networks, optimize agent strategies, and conduct joint evaluation using the local environmental states of other agents.

Benefits of technology

This technology enables agents to adapt to changes in the environment and other agents' policies during multi-agent reinforcement learning, improving learning stability, reducing computational costs, and effectively responding to dynamic network attacks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120639464B_ABST
    Figure CN120639464B_ABST
Patent Text Reader

Abstract

本发明涉及一种基于网络防御智能体的决策方法,涉及网络安全领域,本发明基于网络设备的连接关系构建网络系统拓扑图;利用生成的西格玛规则来在网络系统拓扑图的节点上匹配符合节点属性的网络威胁,从而自动构建出对应网络系统拓扑图中任意节点的攻击动作空间;构建多个攻击智能体和防御智能体;智能体包含决策网络和评价网络;决策网络根据其所观察的局部网络环境状态来给出动作概率,并根据动作概率选择动作;评价网络基于所有攻击智能体观察的联合网络环境状态进行评估得到策略或执行动作的价值;利用强化学习训练攻击智能体和防御智能体,利用训练的防御智能体进行网络防御决策,以应对动态的网络攻击。
Need to check novelty before this filing date? Find Prior Art