Joint Acoustic Visual Processing for Cross-Modal Retrieval

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional automatic speech recognition (ASR) systems require extensive training data and expert knowledge, limiting their applicability to only a few languages, and existing approaches struggle to associate audio and image data without semantic annotation, hindering cross-modal query and retrieval processes.

Innovation Solution

A joint acoustic and visual processing method that uses deep neural networks to associate images with audio signals without requiring semantic annotation, enabling cross-modal similarity processing by configuring image and audio processors to produce comparable numerical representations for similarity evaluation, allowing for query-based retrieval and annotation without human annotation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional ASR systems use large amounts of training data and expert knowledge, then speech recognition accuracy is improved, but the complexity of the system and the cost of accumulating resources increase immensely

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs self-service by automatically learning acoustic units and language structures from untranscribed audio data without requiring manual annotation or expert knowledge. The unsupervised learning algorithms automatically discover repetitions of word-like units, part-of-speech tags, and linguistic patterns, eliminating the need for human annotators and reducing system complexity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical process of manual data annotation and expert knowledge curation with automated computational processes. Machine learning models automatically process raw audio data to extract linguistic features, substituting human expert labor with algorithmic processing that scales more efficiently.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If semantic annotation is used to link images and audio, then the accuracy of cross-modal association is improved, but the requirement for human-annotated training data increases

Engineering Contradiction:
Improvecross-modal association accuracyVSAvoidamount of annotated training data
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system performs self-service by automatically learning to associate images and audio through unsupervised learning from unannotated data. The model discovers correlations between visual and acoustic modalities without requiring human-provided semantic labels, automatically generating its own understanding of cross-modal relationships.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary action by pre-training on large amounts of unannotated image-audio pairs to learn general cross-modal associations before fine-tuning on specific tasks. This preliminary learning from raw data establishes foundational representations that transfer to annotated tasks, reducing the need for extensive labeled training data.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If unsupervised techniques are used to learn acoustic units from untranscribed audio, then the applicability to more languages is improved, but the precision of language learning decreases

Engineering Contradiction:
Improvelanguage applicabilityVSAvoidlanguage learning precision
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The system applies segmentation by breaking down the continuous audio signal into discrete acoustic units through unsupervised clustering. It identifies repetitions and patterns in raw audio to segment word-like units, phonemes, and linguistic structures without pre-defined boundaries, enabling adaptation to any language while maintaining learning precision through iterative refinement.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS10515292B2Joint acoustic and visual processing
Publication Date: 2019.12.24 MASSACHUSETTS INST OF TECH
  • US10515292B2 patent drawing
  • US10515292B2 patent drawing
  • US10515292B2 patent drawing

AI summary

An approach to joint acoustic and visual processing associates images with corresponding audio signals, for example, for the retrievals of images according to voice queries. A set of paired images and audio signals are processed without requiring transcription, segmentation, or annotation of either the images or the audio. This processing of the paired images and audio is used to determine parameters of an image processor and an audio processor, with the outputs of these processors being comparable to determine a similarity across acoustic and visual modalities. In some implementations, the image processor and the audio processor make use of deep neural networks. Further embodiments associate parts of images with corresponding parts of audio signals.