Hierarchical Compositional Networks for Small Training Set Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for visual object recognition and pattern recognition in artificial intelligence, such as computer vision and machine learning, face challenges in recognizing objects with variations and require large training sets, leading to inefficiencies and poor performance.

Innovation Solution

The development of hierarchical compositional networks (HCNs) that utilize convolutional layers and pooling layers to create invariant representations, allowing for effective recognition of objects with variations using smaller training sets by composing parts and introducing variability through pooling.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional deep learning methods are used for visual object recognition, then recognition performance can be achieved, but large training sets are required which increases data collection and processing complexity

Engineering Contradiction:
Improverecognition performanceVSAvoidtraining set size
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent segments the training process into two distinct phases: pre-training on large datasets to learn general features, and fine-tuning on small datasets to specialize in specific tasks. This segmentation allows the system to achieve high recognition performance without requiring the entire training process to use large datasets, thus resolving the contradiction between recognition performance and training set size requirements.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action through the pre-training phase where the model is first trained on large datasets to establish a foundation of general visual features. This preliminary training equips the model with transferable knowledge that reduces the amount of additional training data needed for specific tasks, thereby achieving good performance with smaller fine-tuning datasets.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If models are trained to recognize objects in various positions and settings, then recognition accuracy improves, but the complexity of training data and model structure increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidtraining complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent makes the pre-trained model universal by training it on diverse large datasets that cover multiple object categories, positions, and settings. This universal pre-training creates a model that can handle various recognition tasks without requiring separate specialized training for each scenario, thereby improving recognition accuracy while avoiding the complexity of multiple separate training processes.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent utilizes parameter changes by adjusting the model's learning parameters and architecture during the fine-tuning phase to specialize in specific tasks while maintaining the general capabilities learned during pre-training. This allows the model to achieve high accuracy for specific recognition tasks without retraining the entire model, thus reducing training complexity.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If computers are trained to understand innate properties of objects, then human-like recognition capability is achieved, but traditional training methods require extensive computational resources and time

Engineering Contradiction:
Improvehuman-like recognition capabilityVSAvoidtraining time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent segments training into pre-training and fine-tuning phases, where the bulk of computational resources and time are invested in the pre-training phase on large datasets to establish human-like general recognition capabilities. The fine-tuning phase requires significantly fewer resources, thus achieving human-like capability with reduced overall training time compared to traditional single-phase training methods.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by conducting extensive pre-training before the actual application-specific training. This preliminary action establishes the foundation of human-like recognition capabilities efficiently, and subsequent fine-tuning builds upon this foundation rather than starting from scratch, thereby reducing the total training time required to achieve human-like performance.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11526757B2Systems and methods for deep learning with small training sets
Publication Date: 2022.12.13 INTRINSIC INNOVATION LLC
  • US11526757B2 patent drawing
  • US11526757B2 patent drawing
  • US11526757B2 patent drawing

AI summary

A hierarchical compositional network, representable in Bayesian network form, includes first, second, third, fourth, and fifth parent feature nodes; first, second, and third pool nodes; first, second, and third weight nodes; and first, second, third, fourth, and fifth child feature nodes.