Visual Question Answering Model Using Speech-Visual Embeddings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing visual question answering models do not fully utilize the information in coordinated text-video representations and fail to leverage relationships between different modal spaces, such as speech and vision, especially when trained on unlabeled videos without annotation.

Innovation Solution

A system that learns a shared embedding space using speech-visual correspondence on unlabeled videos, generating additional embeddings for question, video, and answer, enabling the training of a visual question answering model that can perform tasks like action recognition, object recognition, and captioning without requiring labeled data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If visual question answering models are trained using traditional methods with labeled data, then they can achieve basic VQA performance, but they fail to fully utilize information in coordinated text-video representations and do not leverage relationships between different modal spaces

Engineering Contradiction:
Improveinformation utilization in text-video representationsVSAvoidVQA performance
Core Design Contradiction:
Loss of informationVSReliability

Solution Approach 1:

The patent merges speech and vision modalities into a shared embedding space, allowing the model to leverage relationships between different modal spaces. By embedding both speech representations and visual representations in the same vector space, the model can utilize coordinated text-video representations more effectively, resolving the contradiction between information utilization and performance reliability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The shared embedding space serves multiple functions: it enables speech-visual correspondence learning, supports visual question answering, and facilitates action recognition. This multi-functional approach allows the same embedding space to be used for different tasks, improving both information utilization and VQA performance simultaneously.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of manufacture

If models are trained on unlabeled videos without annotation, then data availability increases and annotation costs decrease, but existing models cannot effectively learn from these unlabeled videos

Engineering Contradiction:
Improvemodel training ease with unlabeled dataVSAvoidlearning effectiveness from unlabeled videos
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The model performs self-supervised learning by learning speech-visual correspondence from unlabeled videos without requiring external annotations. The system uses the inherent structure and relationships within the unlabeled data itself to train the embedding space, making the training process self-service and eliminating the need for manual labeling while maintaining learning effectiveness.

Inventive Principle:
Principle #25Self-service

3Loss of information

If a shared embedding space is learned using speech-visual correspondence on unlabeled videos, then the model can leverage freely available data and improve information utilization, but the complexity of learning multiple embeddings in coordinated spaces increases

Engineering Contradiction:
Improveutilization of unlabeled video informationVSAvoidembedding learning complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The shared embedding space acts as an intermediary that bridges speech and vision modalities. By introducing this intermediate representation space, the model can learn complex relationships between different modalities without directly comparing speech and visual features, thereby managing the complexity of learning multiple embeddings while effectively utilizing unlabeled video information.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS20220067546A1Visual question answering using model trained on unlabeled videos
Publication Date: 2022.03.03 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US20220067546A1 patent drawing
  • US20220067546A1 patent drawing
  • US20220067546A1 patent drawing

AI summary

An example system includes a processor to learn a shared embedding space on unlabeled videos using speech visual correspondence. The processor can learn a number of additional embeddings including a question plus video embedding and an answer embedding using the shared embedding space to generate a trained visual question answering model. The processor can execute a visual question answering based on the trained visual question answering model.