Excavator Control Using Rewarded Position Curves

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing excavator control systems require complex and time-consuming manual adjustments of control algorithms for combined operations like leveling ground or brushing slopes, which are labor-intensive and costly.

Innovation Solution

A working machine control method utilizing reinforcement learning to automatically adjust control algorithms based on a state-behavior decision model, determining a reward value from the coincidence between actual and target position curves of the working portion, allowing the excavator to perform construction tasks efficiently without requiring precise control models for each state.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If traditional control algorithms are used for excavator combined operations, then control accuracy can be achieved, but the complexity of adjusting control algorithms increases significantly and requires extensive manual work

Engineering Contradiction:
Improvecontrol accuracyVSAvoidcontrol algorithm complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system uses reinforcement learning to enable the excavator to automatically learn and optimize control strategies through trial-and-error training. The agent independently adjusts control parameters based on reward feedback from position curve coincidence, eliminating the need for manual control algorithm adjustment by engineers while achieving high control accuracy in combined operations

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The reinforcement learning model dynamically changes control parameters during training and operation. By adjusting the state space, action space, and reward function parameters, the system adapts to different operation conditions (leveling, slope brushing) without requiring separate manual tuning for each scenario, reducing overall system complexity

Inventive Principle:
Principle #35Parameter changes

2Manufacturing precision

If traditional control algorithms are manually adjusted for each operation state, then desired control accuracy can be achieved, but the time and labor cost for adjustment increases significantly

Engineering Contradiction:
Improvecontrol accuracyVSAvoidadjustment time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary automated training offline to pre-adjust control algorithms for various operation states. During actual operation, the pre-trained model directly provides optimal control decisions without requiring real-time manual adjustment. The reinforcement learning agent learns optimal policies in advance through simulated training, significantly reducing on-site adjustment time and labor costs

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The reinforcement learning system automatically performs all adjustment tasks that previously required manual engineer intervention. Through self-learning from state-behavior-reward tuples, the system independently optimizes control parameters for different operation conditions, eliminating time-consuming manual adjustment processes while maintaining high control accuracy

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12540455B2Working machine control method using target position curve and reward model, working machine control device and working machine
Publication Date: 2026.02.03 SHANGHAI SANY HEAVY IND
  • US12540455B2 patent drawing
  • US12540455B2 patent drawing
  • US12540455B2 patent drawing

AI summary

Disclosed are a working machine control method and a device, and a working machine. The method includes: obtaining a current working state of a working machine; determining a current decision behavior of the work machine based on the current work state and a state-behavior decision model; and controlling, based on a control signal corresponding to the current decision behavior, the work machine to perform construction work. The state-behavior decision model is based on a sample working state, a sample decision behavior, and a reward value corresponding to the sample decision behavior. The reward value is determined based on an actual position curve and target position curve; the actual position curve is determined based on the sample decision behavior. The method, device and working machine reduce the adjusting workload of engineers, shorten the adjusting time, reduce the adjusting cost, and improve the intelligent construction level of the working machine.