Behavior Learning System for Similarity-Controlled Attack Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional reinforcement learning techniques for generating attack data in false information injection attacks fail to control neural networks to acquire actions based on the closeness between the environment's properties with and without the attack, limiting the ability to generate data that is difficult or easy to detect as abnormal.

Innovation Solution

An action learning system that includes a first training unit to train a second neural network to calculate similarity degrees between environmental data with and without applied actions, and a second training unit to train a first neural network using a reward value based on these similarity degrees to determine optimal actions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If only influence degree is used as reward in reinforcement learning, then the neural network can acquire actions that maximize attack impact, but the ability to control the similarity between environmental properties with and without attacks is lost

Engineering Contradiction:
Improveattack data generation capabilityVSAvoidcontrol over environmental property similarity
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The reward function is segmented into two independent components: influence degree (attack impact) and similarity degree (environmental property closeness). This allows the neural network to be trained to optimize one aspect while controlling the other, resolving the contradiction between maximizing attack impact and controlling property similarity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a new parameter (similarity degree) to the reward function alongside the existing influence degree parameter. By changing the reward structure to include both parameters, the system gains control over the similarity between environmental properties while maintaining attack effectiveness.

Inventive Principle:
Principle #35Parameter changes

2Object-affected harmful factors

If attack data is generated to be difficult to detect, then detection capability is reduced, but the ability to control the similarity to normal environmental properties is compromised

Engineering Contradiction:
Improvedetectability of attackVSAvoidcontrol over attack data properties
Core Design Contradiction:
Object-affected harmful factorsVSManufacturing precision

Solution Approach 1:

The similarity degree calculation provides feedback to the reinforcement learning process, allowing the neural network to adjust its actions to maintain environmental property similarity while achieving detection evasion. This feedback mechanism enables precise control over attack data properties.

Inventive Principle:
Principle #23Feedback

3Object-affected harmful factors

If attack data is generated to be easy to detect, then detection capability is enhanced, but the ability to control the similarity to normal environmental properties is reduced

Engineering Contradiction:
Improvedetectability of attackVSAvoidcontrol over attack data properties
Core Design Contradiction:
Object-affected harmful factorsVSManufacturing precision

Solution Approach 1:

The system dynamically adjusts the balance between influence degree and similarity degree based on detection requirements. The reward function can be configured to prioritize either component, allowing the system to adaptively generate attack data with desired detectability characteristics while maintaining control over property similarity.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12412075B2Behavior learning system, behavior learning method and program
Publication Date: 2025.09.09 NT T INC
  • US12412075B2 patent drawing
  • US12412075B2 patent drawing
  • US12412075B2 patent drawing

AI summary

An action learning system includes a memory, and a processor configured to train, based on first data indicating a property of an environment in which data is collected from multiple devices and to which an action determined by a first neural network according to a state of the environment is applied and second data indicating a property of the environment to which the action is not applied, a second neural network that calculates a similarity degree between distributions of the first data and the second data, and train, after the second neural network is trained, the first neural network that determines an action according to the state of the environment, by reinforcement learning including, in a reward, a value that changes based on a relationship between a similarity degree and a parameter set by a user, the similarity degree being calculated by the second neural network based.