Reinforcement Learning for Servo Motor Reactive Current Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional servo motor controllers require complex setting and adjustment of base and clamp velocities to manage reactive current, which can lead to voltage saturation issues, especially during high-velocity rotations, and do not adapt effectively to changes over time due to aging.

Innovation Solution

A machine learning device that performs reinforcement learning to calculate an appropriate reactive current command for servo motors without pre-setting base and clamp velocities, using state information and feedback to optimize the command and avoid voltage saturation through a reward-based system.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If base velocity and clamp velocity are manually set to control reactive current, then voltage saturation can be avoided, but the setting operation becomes complex and requires frequent adjustments due to aging

Engineering Contradiction:
Improvevoltage saturation preventionVSAvoidsetting operation complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The servo motor controller automatically determines base velocity and clamp velocity through internal calculations based on motor parameters and operating conditions, eliminating the need for manual setting and adjustment. The system serves itself by computing optimal velocity thresholds dynamically

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The base velocity and clamp velocity are not fixed values but are dynamically determined based on motor parameters, load conditions, and operating state. The system adapts these velocity thresholds in real-time rather than using static pre-set values

Inventive Principle:
Principle #15Dynamics

2Reliability

If reactive current is increased to reduce counter-electromotive force in high-velocity region, then stable rotation control is achieved, but heat generation due to reactive current increases

Engineering Contradiction:
Improverotation control stabilityVSAvoidheat generation
Core Design Contradiction:
ReliabilityVSTemperature

Solution Approach 1:

The system supplies reactive current to the d-phase only when necessary (in high-velocity regions where voltage saturation occurs), rather than continuously. By controlling the magnitude and timing of reactive current injection, the system achieves stable rotation control while minimizing unnecessary heat generation from continuous reactive current flow

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system dynamically adjusts the d-phase current command based on the operating velocity and load conditions. By changing the reactive current parameter adaptively rather than maintaining a fixed value, the system optimizes the balance between rotation stability and heat generation

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10684594B2Machine learning device, servo motor controller, servo motor control system, and machine learning method
Publication Date: 2020.06.16 FANUC LTD
  • US10684594B2 patent drawing
  • US10684594B2 patent drawing
  • US10684594B2 patent drawing

AI summary

A machine learning device performs machine learning with respect to a servo motor controller that converts a three-phase current to a two-phase current of the d- and q-phase. The machine learning device includes: a state information acquisition unit configured to acquire, from the servo motor controller, state information including velocity or a velocity command, reactive current, and an effective current command and effective current or a voltage command; an action information output unit configured to output action information including a reactive current command to the servo motor controller; a reward output unit configured to output a value of a reward of reinforcement learning based on the voltage command or the effective current command and the effective current; and a value function updating unit configured to update a value function on the basis of the output value of the reward, the state information, and the action information.