一种捕捉环境中可控因素的表示学习方法及系统

By introducing the concept of controllable factors and mutual information measurement, a representation learning method was designed to solve the problem of noise interference in reinforcement learning and improve the robustness and accuracy of policy training.

CN117688983BActive Publication Date: 2026-07-17INST OF COMPUTING TECH CHINESE ACAD OF SCI

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INST OF COMPUTING TECH CHINESE ACAD OF SCI
Filing Date
2022-08-23
Publication Date
2026-07-17

AI Technical Summary

Technical Problem

Existing reinforcement learning algorithms struggle to effectively distinguish predictable noise in the environment when processing high-dimensional observation data, causing task-related information to be squeezed out in low-dimensional representations and affecting policy training performance.

Method used

By introducing the concept of controllable factors, a representation learning method is designed. The method maximizes the content of controllable factors using mutual information metric and loss function. Convolutional neural networks and multi-layer fully connected networks are used to construct the encoder to filter noise and improve the robustness of representation learning.

Benefits of technology

Effective filtering of predictable noise in the environment improves the robustness of representation learning and enhances the accuracy and efficiency of policy training.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117688983B_ABST
    Figure CN117688983B_ABST
Patent Text Reader

Abstract

本发明提出一种捕捉环境中可控因素的表示学习方法和系统,包括:智能体采集在当前所处环境的观测图像,通过卷积神经网络对该观测图像进行编码,得到当前时刻t该观测图像的表示;统计该当前时刻t该观测图像的表示、t时刻到t+k‑1时刻策略所采取的动作序列和第t+k时刻该观测图像的表示,三者之间的互信息作为可控因素的度量;基于该度量构建损失函数,以最大化该度量,基于该度量最大时对应的时刻t该观测图像的表示,执行学习策略,得到目标动作,该智能体执行该目标动作与该环境产生交互。本发明通过捕捉环境中的可控因素,能有效过滤其他可预测的噪声,因此在复杂环境上具备更好的鲁棒性。
Need to check novelty before this filing date? Find Prior Art