Deep Compositional Framework for Zero-Shot Language Navigation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current machine learning techniques lack a reasonably fast learning rate for achieving human-like zero-shot learning ability, which is essential for tasks like object recognition and navigation, and fail to ground language in vision effectively.

Innovation Solution

A deep compositional framework that integrates vision and language, allowing an agent to learn from scratch through gradient descent, with a modular architecture that enables zero-shot learning by grounding language in vision and transferring knowledge from recognition to navigation tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If current machine learning techniques are used, then the system can process language tasks, but the learning rate for zero-shot learning is unacceptably slow

Engineering Contradiction:
Improvelearning rateVSAvoidzero-shot learning ability
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The language processing system is segmented into distinct modules: a language module that processes linguistic input and a vision module that processes visual input. This segmentation allows each module to specialize in its domain while enabling efficient knowledge transfer for zero-shot learning, resolving the contradiction between fast learning rate and reliable generalization ability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary action by pre-training on language-instruction tasks before deploying to navigation tasks. This preliminary training establishes foundational language understanding and visual grounding that enables rapid zero-shot transfer to new tasks, achieving both fast learning rate and reliable performance.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If a sophisticated language system is developed, then human-level intelligence can be achieved, but the system complexity increases significantly

Engineering Contradiction:
Improvelanguage understanding capabilityVSAvoidsystem architecture
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system divides complex language processing into modular components including syntax processing, semantics processing, and visual grounding modules. This segmentation reduces overall system complexity while maintaining sophisticated language understanding capabilities through specialized sub-components.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The language module serves multiple functions: processing natural language instructions, generating navigation commands, and grounding language in visual concepts. This multi-functionality achieves versatile language understanding without proportionally increasing system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If language semantics are grounded in perception experience, then knowledge can be transferred from task to task, but the training process becomes more complex

Engineering Contradiction:
Improveknowledge transfer abilityVSAvoidtraining process
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs preliminary visual grounding during the language training phase, establishing associations between language semantics and visual perception before task deployment. This preliminary grounding enables automatic knowledge transfer to new tasks without adding complexity to the training process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The language module automatically grounds itself in visual perception through self-supervised learning from the environment, eliminating the need for manual annotation or complex external training procedures. This self-service approach enables knowledge transfer while keeping the training process simple.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS10366166B2Deep compositional frameworks for human-like language acquisition in virtual environments
Publication Date: 2019.07.30 BAIDU USA LLC
  • US10366166B2 patent drawing
  • US10366166B2 patent drawing
  • US10366166B2 patent drawing

AI summary

Described herein are systems and methods for human-like language acquisition in a compositional framework to implement object recognition or navigation tasks. Embodiments include a method for a model to learn the input language in a grounded and compositional manner, such that after training the model is able to correctly execute zero-shot commands, which have either combination of words in the command never appeared before, and/or new object concepts learned from another task but never learned from navigation settings. In embodiments, a framework is trained end-to-end to learn simultaneously the visual representations of the environment, the syntax and semantics of the language, and outputs actions via an action module. In embodiments, the zero-shot learning capability of a framework results from its compositionality and modularity with parameter tying.