Feedback-Graph Action Selection for Dynamic Online Optimization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing online optimization methods fail to determine the best action dynamically due to fixed comparisons with static regret, leading to suboptimal decision-making in uncertain environments.

Innovation Solution

An information processing device and method that utilizes a feedback graph to determine a probability distribution for action selection, observes losses, and updates weights based on these losses to minimize dynamic regret, incorporating algorithms for strongly and weakly observable feedback graphs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If static regret is used to determine the target action, then the comparison action is fixed at each round, but the determined action may not be the best action or similar to the best action

Engineering Contradiction:
Improveaction determination accuracyVSAvoidadaptability to changing optimal actions
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent transforms the static regret framework into a dynamic one by introducing time-varying probability distributions over actions. Instead of fixing a single target action, the system maintains a distribution that evolves based on observed losses and feedback graphs, allowing the comparison target to adapt dynamically to changing conditions while improving action determination accuracy.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes the parameter from a fixed target action to a probability distribution over actions. By updating the probability distribution based on observed losses and feedback graph structures, the system can adapt to changing optimal actions without being constrained by a predetermined fixed target, thus resolving the contradiction between reliability and adaptability.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If a fixed action is assumed as target for static regret, then the calculation is simple, but the determined action is not necessarily optimal in uncertain environments

Engineering Contradiction:
Improvedecision-making efficiencyVSAvoidoptimality of determined action
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent introduces feedback mechanisms where observed losses from actions are fed back into the system to update probability distributions. This feedback loop allows the system to learn from actual outcomes and adjust future action selections, improving both the reliability of optimal action determination and maintaining decision-making efficiency through iterative refinement.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent performs preliminary exploration by maintaining probability distributions over multiple actions rather than committing to a single fixed action. This preliminary diversification of action choices allows the system to gather information about different actions' performances before converging on the optimal choice, improving reliability without significantly reducing productivity.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250307730A1Information processing device, control method, and storage medium
Publication Date: 2025.10.02 NEC CORP
  • US20250307730A1 patent drawing
  • US20250307730A1 patent drawing
  • US20250307730A1 patent drawing

AI summary

The information processing device 1X mainly includes an acquisition means 15Xa, a first determination means 15Xb, a selection means 15Xc, an observation means 15Xd, and a second determination means 15Xe. The acquisition means 15Xa acquires a feedback graph representing an online optimization problem. The first determination means 15Xb determines, based on the feedback graph, a probability distribution for selecting an action to be taken from action candidates. The selection means 15Xc selects the action based on the probability distribution. The observation means 15Xd observes a loss based on the action. The second determination means 15Xe determines, based on the observed loss, a weight for determining the probability distribution.