Grounded Language Models for Human Activity Forecasting in Task Planning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current robotic systems lack the ability to proactively infer human intentions and predict future actions, leading to inadequate task and motion planning in human-centric environments, as they rely on reactive collision avoidance and short-term trajectory predictions without considering long-term human motion or environmental context.
Innovation Solution
A method utilizing geometrically grounded large language models to convert sensor data into natural language narrations, which are then combined with sequence text to predict human activities and correlate them with locations in a semantic map, enabling the determination of relevancy scores and control signals for device movement to minimize human-robot interactions and optimize task planning.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If behavioral models are trained to predict human motion from past activities, then prediction capability is improved, but large amounts of data are required and models are constrained to specific data formats and modalities
Solution Approach 1:
The patent replaces traditional behavior cloning approaches with large language models that use natural language descriptions of activities instead of requiring extensive labeled motion data. The LLM processes textual activity descriptions to predict future human actions, substituting the data-intensive mechanical training process with a language-based reasoning approach that generalizes better across different scenarios
Solution Approach 2:
The patent changes the input parameter format from structured motion data and trajectories to natural language activity descriptions. This parameter transformation allows the system to leverage pre-trained language models that can reason about human intentions without requiring retraining on specific motion datasets, thereby reducing data requirements while maintaining prediction reliability
2Reliability
If robots use reactive collision avoidance, then immediate safety is improved, but the ability to proactively infer human intentions and plan tasks is degraded
Solution Approach 1:
The patent implements preliminary action by using large language models to predict future human activities and intentions before they occur. The system analyzes current activity descriptions and generates predictions about upcoming actions, allowing the robot to proactively plan tasks and adjust behavior in advance rather than merely reacting to collisions or immediate threats
Solution Approach 2:
The patent introduces natural language activity descriptions as an intermediary between raw sensor data and robot decision-making. This linguistic representation layer enables the system to infer human intentions and predict future actions while maintaining safety, bridging the gap between reactive collision avoidance and proactive task planning through interpretable activity reasoning
3Speed
If predictions are made without environmental context, then processing speed is improved, but the ability to localize predictions in physical environment is degraded
Solution Approach 1:
The patent merges the large language model's prediction capabilities with environmental context from semantic maps and sensor data. The system combines textual activity descriptions with spatial information about the physical environment, allowing predictions to be localized to specific locations and objects while maintaining processing efficiency through the LLM's inherent understanding of spatial relationships in natural language
Data Source
AI summary
A method includes receiving data from one or more sensors, detecting, from the received data, one or more activities of a person in a space, converting the detected one or more activities into a natural language narration; extracting one or more items from a map corresponding to the space; determining a relevancy score for each of the extracted one or more items based on an output of a language model that receives as input a combination of the narration with a binding sequence; correlating the extracted one or more items with one or more locations on the map based on the determined relevancy for each of the extracted one or more items; and outputting a control signal for controlling a movement of one more devices based on the correlation of the extracted one or more items with the one or more locations on the map.


