Soft Layer Ordering in Deep Multitask Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing deep multitask learning approaches are limited by the assumption of parallel ordering of layers, which restricts the sharing of learned transformations across tasks and limits the scope of multitask learning to only closely-related tasks, failing to capture diverse structural regularities found in complex real-world tasks.

Innovation Solution

The introduction of soft ordering of layers, where shared layers are applied in different orders for different tasks, allowing the model to learn how to apply shared layers in various ways at different depths, enabling more flexible integration of information across tasks and enabling the model to learn generalizable modules that can be assembled in novel ways for future unseen tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If parallel ordering of layers is assumed in deep multitask learning, then the model structure is simplified and training is easier, but the ability to capture diverse structural regularities across unrelated tasks is limited

Engineering Contradiction:
Improvemodel structure complexityVSAvoidability to capture diverse structural regularities
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent applies dynamics by making the layer ordering flexible and task-specific rather than fixed and parallel. Each task can have its own ordering of shared layers, allowing the model to adapt the structural configuration dynamically based on task requirements while still benefiting from parameter sharing across tasks.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent implements local quality by allowing different tasks to have different layer orderings in specific regions of the network. While the shared layers are common across tasks, their ordering can be customized locally for each task, enabling tailored feature extraction sequences for different task types without increasing overall model complexity.

Inventive Principle:
Principle #3Local quality

2Reliability

If shared layers are applied in the same order for all tasks, then the training process is more stable and converges faster, but the model cannot effectively learn from superficially unrelated tasks

Engineering Contradiction:
Improvetraining stabilityVSAvoidability to learn from unrelated tasks
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent resolves this contradiction by introducing dynamic task-specific ordering of shared layers. The model maintains stable training through consistent parameter sharing while allowing the sequence of layer application to vary by task, enabling effective learning from unrelated tasks through customized feature extraction orderings.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies segmentation by dividing the shared layers into task-specific sequences. Each task receives a segmented or permuted version of the shared layers, allowing independent optimization of layer ordering for each task while maintaining the shared nature of the underlying parameters.

Inventive Principle:
Principle #1Segmentation

3Ease of manufacture

If fixed ordering of shared layers is used, then the model is easier to implement and train, but it fails to exploit the full potential of deep multitask learning for diverse tasks

Engineering Contradiction:
Improveimplementation easeVSAvoidlearning effectiveness
Core Design Contradiction:
Ease of manufactureVSProductivity

Solution Approach 1:

The patent balances implementation ease with learning effectiveness by introducing dynamic task-specific layer ordering. The implementation remains relatively simple through parameter sharing, while the dynamic ordering capability enables the model to exploit diverse structural regularities across tasks, significantly improving learning effectiveness.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies universality by creating a single shared set of layers that can be universally applied across multiple tasks with different orderings. This multi-functional approach allows the same layers to serve different purposes in different task sequences, maximizing learning effectiveness without requiring separate models for each task.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11250314B2Beyond shared hierarchies: deep multitask learning through soft layer ordering
Publication Date: 2022.02.15 COGNIZANT TECHNOLOGY SOLUTIONS US CORP
  • US11250314B2 patent drawing
  • US11250314B2 patent drawing
  • US11250314B2 patent drawing

AI summary

The technology disclosed identifies parallel ordering of shared layers as a common assumption underlying existing deep multitask learning (MTL) approaches. This assumption restricts the kinds of shared structure that can be learned between tasks. The technology disclosed demonstrates how direct approaches to removing this assumption can ease the integration of information across plentiful and diverse tasks. The technology disclosed introduces soft ordering as a method for learning how to apply layers in different ways at different depths for different tasks, while simultaneously learning the layers themselves. Soft ordering outperforms parallel ordering methods as well as single-task learning across a suite of domains. Results show that deep MTL can be improved while generating a compact set of multipurpose functional primitives, thus aligning more closely with our understanding of complex real-world processes.