Knowledge Module for Interpretable Multi-Modal Deep Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional deep learning models are limited in their ability to process multi-modal inputs, adapt to new domains, and provide interpretable decision-making processes, relying heavily on statistical patterns and extensive training data, which hinders their ability to understand deeper knowledge and interact naturally with users.

Innovation Solution

The development of systems and methods that generate and utilize knowledge modules to integrate multi-modal embeddings, allowing for dynamic updates and external attention, enabling the model to learn new knowledge without retraining the main model and providing interpretability through knowledge attention and causal inference.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional scalable models are used to improve accuracy in perception tasks, then model performance is improved, but the models remain limited to statistical learning patterns in a black-box manner and cannot access other modalities or provide interpretable decision-making

Engineering Contradiction:
Improveaccuracy in perception tasksVSAvoidability to access other modalities and provide interpretable decision-making
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system segments the deep learning model into distinct components: a main model for statistical learning and a knowledge module for structured knowledge storage. This segmentation allows the model to maintain high accuracy through the main model while gaining interpretability and multi-modality access through the separate knowledge module, resolving the contradiction between accuracy and versatility.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The knowledge module acts as an intermediary between the main model and external knowledge sources. It receives knowledge inputs from multiple modalities, processes them through representation models and grounding types, and provides structured knowledge embeddings to the main model. This intermediary enables the system to access other modalities and provide interpretable decision-making without compromising the main model's accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If extensive supervised learning is used to train conventional models, then model performance is improved, but the need for extensive labeled data and training iterations increases

Engineering Contradiction:
Improvemodel performanceVSAvoidextensive training iterations and labeled data requirements
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by pre-processing knowledge inputs into structured knowledge embeddings before they are needed by the main model. The knowledge module pre-computes representations and grounds knowledge units in advance, creating a ready-to-use knowledge base that reduces the need for extensive training iterations when the model needs to adapt to new domains or tasks.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of re-training the entire main model with extensive labeled data, the system creates copies or embeddings of knowledge from various modalities and stores them in the knowledge module. This copying approach allows the model to leverage pre-processed knowledge representations, significantly reducing training time and labeled data requirements while maintaining performance.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If conventional models are updated to incorporate new domain knowledge, then the models can adapt to new domains, but significant additional training on large training datasets is required

Engineering Contradiction:
Improveability to adapt to new domainsVSAvoidadditional training data requirements
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The system extracts domain-specific knowledge from knowledge inputs and stores it separately in the knowledge module, taking it out from the main model's training data requirements. When adapting to new domains, only the knowledge module needs to be updated with new knowledge embeddings, while the main model remains unchanged. This extraction approach enables domain adaptation without requiring additional training data for the main model.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The knowledge module is designed to be dynamic and easily updatable. New domain knowledge can be incorporated by simply adding new knowledge inputs to the knowledge module, which then processes them into embeddings. This dynamic structure allows the system to adapt to new domains flexibly without the rigid constraint of re-training the entire model on large datasets.

Inventive Principle:
Principle #15Dynamics

4Measurement precision

If large-scale pre-trained models are used to improve accuracy, then performance in various perception tasks is improved, but the gap between model performance and human cognition remains large

Engineering Contradiction:
Improveperformance in perception tasksVSAvoidability to understand deeper level knowledge like concepts, relations, and commonsense
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system creates a composite architecture combining the main model (for statistical learning and accuracy) with the knowledge module (for structured knowledge representation). This composite structure integrates two different approaches: the computational power of deep learning and the structured reasoning capabilities of knowledge graphs. The knowledge module stores and processes deeper level knowledge such as concepts, relations, and commonsense in a structured manner, complementing the main model's perception capabilities and bridging the gap toward human-like cognition.

Inventive Principle:
Principle #40Composite materials

Data Source

PatentUS20230229960A1Systems and methods for facilitating integrative, extensible, composable, and interpretable deep learning
Publication Date: 2023.07.20 MICROSOFT TECHNOLOGY LICENSING LLC
  • US20230229960A1 patent drawing
  • US20230229960A1 patent drawing
  • US20230229960A1 patent drawing

AI summary

Some disclosed systems are configured to obtain a knowledge module configured to receive one or more knowledge inputs corresponding to one or more different modalities and generate a set of knowledge embeddings to be integrated with a set of multi-modal embeddings generated by a multi-modal main model. The systems receive a knowledge input at the knowledge module, identify a knowledge type associated with the knowledge input, and extract a knowledge unit from the knowledge input. The systems select a representation model that corresponds to the knowledge type and select a grounding type configured to ground the at least one knowledge unit into the representation model. The systems then ground the knowledge unit into the representation model according to the grounding type.