Behavior Learning System for Similarity-Controlled Attack Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional reinforcement learning techniques for generating attack data in false information injection attacks fail to control neural networks to acquire actions based on the closeness between the environment's properties with and without the attack, limiting the ability to generate data that is difficult or easy to detect as abnormal.
Innovation Solution
An action learning system that includes a first training unit to train a second neural network to calculate similarity degrees between environmental data with and without applied actions, and a second training unit to train a first neural network using a reward value based on these similarity degrees to determine optimal actions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If only influence degree is used as reward in reinforcement learning, then the neural network can acquire actions that maximize attack impact, but the ability to control the similarity between environmental properties with and without attacks is lost
Solution Approach 1:
The reward function is segmented into two independent components: influence degree (attack impact) and similarity degree (environmental property closeness). This allows the neural network to be trained to optimize one aspect while controlling the other, resolving the contradiction between maximizing attack impact and controlling property similarity.
Solution Approach 2:
The patent introduces a new parameter (similarity degree) to the reward function alongside the existing influence degree parameter. By changing the reward structure to include both parameters, the system gains control over the similarity between environmental properties while maintaining attack effectiveness.
2Object-affected harmful factors
If attack data is generated to be difficult to detect, then detection capability is reduced, but the ability to control the similarity to normal environmental properties is compromised
Solution Approach 1:
The similarity degree calculation provides feedback to the reinforcement learning process, allowing the neural network to adjust its actions to maintain environmental property similarity while achieving detection evasion. This feedback mechanism enables precise control over attack data properties.
3Object-affected harmful factors
If attack data is generated to be easy to detect, then detection capability is enhanced, but the ability to control the similarity to normal environmental properties is reduced
Solution Approach 1:
The system dynamically adjusts the balance between influence degree and similarity degree based on detection requirements. The reward function can be configured to prioritize either component, allowing the system to adaptively generate attack data with desired detectability characteristics while maintaining control over property similarity.
Data Source
AI summary
An action learning system includes a memory, and a processor configured to train, based on first data indicating a property of an environment in which data is collected from multiple devices and to which an action determined by a first neural network according to a state of the environment is applied and second data indicating a property of the environment to which the action is not applied, a second neural network that calculates a similarity degree between distributions of the first data and the second data, and train, after the second neural network is trained, the first neural network that determines an action according to the state of the environment, by reinforcement learning including, in a reward, a value that changes based on a relationship between a similarity degree and a parameter set by a user, the similarity degree being calculated by the second neural network based.


