Reinforcement-Learning Task Assignment for Energy-Limited Work Machines
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current systems for managing task assignments of work machines in construction, mining, and excavating operations are inefficient, particularly for greenhouse gas-free machines, which require frequent recharging and have limited energy storage, and do not account for the unique challenges of these operations, such as varying site conditions and machine operating conditions.
Innovation Solution
A computer-implemented method using a reinforcement-learning model that receives state data from work machines and site conditions to optimize task assignments, predicting performance and energy consumption, and selecting tasks to maximize efficiency and energy use, while ensuring the machines operate within their capacity and recharge needs.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual dispatch is used to manage work machines, then operational flexibility is maintained, but productivity and efficiency are insufficient
Solution Approach 1:
The system enables autonomous work machines to self-manage their task assignments by equipping them with onboard processors that locally execute machine learning models. Each machine independently evaluates its own state (battery charge, location, capabilities) and autonomously selects optimal tasks from available opportunities, eliminating the need for centralized manual dispatch while maximizing productivity.
Solution Approach 2:
The system transitions from static, rule-based dispatch parameters to dynamic, adaptive parameters by implementing reinforcement learning models that continuously learn from operational data. The machine learning models process real-time state data (battery state-of-charge, location, task requirements) and adaptively determine optimal task assignments, transforming the dispatch management approach from manual to intelligent automated decision-making.
2Object-affected harmful factors
If greenhouse gas-free machines with limited energy storage are used, then environmental sustainability is improved, but operational duration and productivity are reduced
Solution Approach 1:
The system performs preliminary evaluation of battery state-of-charge levels before task assignment. The reinforcement learning model predicts future energy requirements based on current state and task characteristics, proactively identifying when recharging is needed. This allows the system to schedule recharging activities in advance, ensuring machines maintain sufficient energy for productive operations while maximizing the utilization of green energy sources.
Solution Approach 2:
The system implements dynamic task assignment that adapts to real-time battery state-of-charge levels. As battery charge decreases, the reinforcement learning model dynamically adjusts task selection to favor lower-energy-consuming tasks or schedules recharging during low-productivity periods. This dynamic adaptation optimizes the balance between environmental sustainability and operational duration, allowing green machines to operate efficiently throughout their energy capacity.
3Use of energy by moving object
If frequent recharging is required for green machines, then energy storage constraints are addressed, but idle time increases and productivity decreases
Solution Approach 1:
The system implements continuous feedback loops where work machines report their state-of-charge levels, location, and operational status to the reinforcement learning model in real-time. The model processes this feedback and dynamically adjusts task assignments and recharging schedules. This feedback mechanism enables the system to optimize energy management by identifying optimal recharging moments that minimize idle time, such as scheduling charges during natural breaks in the work cycle or when machines are positioned near charging infrastructure.
Solution Approach 2:
The reinforcement learning model works to maintain continuous productive operation by strategically scheduling recharging activities during periods of naturally lower productivity demand. By predicting work cycle patterns and energy consumption rates, the system plans recharging to occur during transitions between tasks or during low-priority periods, thereby minimizing interruptions to the overall productive action and reducing idle time.
4Productivity
If automated dispatch based on restrictions is implemented, then dispatch efficiency is improved, but continuous controller intervention is still required
Solution Approach 1:
The system achieves full automation by enabling work machines to autonomously evaluate their own state and independently select optimal tasks using onboard reinforcement learning models. Each machine self-determines its next action based on real-time data about its battery charge, location, capabilities, and available tasks, completely eliminating the need for continuous controller intervention while maintaining high dispatch efficiency.
Solution Approach 2:
The system replaces the mechanical control system requiring human controller intervention with an intelligent software-based reinforcement learning model. The machine learning algorithms process complex multi-factor decision-making (energy constraints, task requirements, machine state) that would be difficult for controllers to optimize manually, substituting human cognitive processing with automated computational intelligence that operates continuously without fatigue or error.
Data Source
AI summary
Systems and methods are disclosed for managing task assignments for a plurality of work machines at a site. An assignment engine may: receive first state data for a work machine including historical data, operating condition, and location data, and second state data for the site including characteristic data for materials and a plurality of available tasks; predict performance data and energy consumption data of the work machine for a task; select a task for the work machine by inputting first state data and second state data into a trained reinforcement-learning model, wherein: the model has been trained to learn an assignment policy that optimizes a reward function such that the learned policy selects a task for at least one work machine from the plurality of tasks available at the site; and cause the at least one work machine to be operated according to the at least one task assignment.


