Audio Input Analysis for Speech and Music Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current audio recognition systems for speech and music operate independently, leading to a negative user experience as users must select different applications for speech or music recognition, and inefficient use of computational resources due to differing processing costs.

Innovation Solution

Integrating speech and music recognition services into a single application that determines the type of audio input to provide appropriate services automatically, conserving resources by processing audio input for music recognition only when necessary, using acoustic fingerprints for music identification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If speech and music recognition services operate independently, then each service can be optimized for its specific function, but users must select different applications and computational resources are wasted processing audio for both services simultaneously

Engineering Contradiction:
Improvecomputational efficiencyVSAvoiduser convenience
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent combines speech recognition and music recognition services into a single integrated application. The system includes a music recognition module and a speech recognition module that share common components such as audio input processing, endpoint detection, and result presentation interfaces. This merging allows the system to handle both types of audio input through one unified interface, improving user convenience while enabling intelligent routing to avoid processing waste.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system dynamically determines the type of audio input (speech or music) and routes it to the appropriate recognition module. The integration layer analyzes incoming audio and activates only the necessary processing pipeline - either music recognition or speech recognition - based on the detected input type. This dynamic adaptation resolves the contradiction by maintaining service optimization while preventing redundant computational processing.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If the system processes audio input for both speech and music recognition simultaneously, then comprehensive service coverage is provided, but computational resources are inefficiently utilized due to differing processing costs

Engineering Contradiction:
Improveservice coverageVSAvoidcomputational resource consumption
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The system segments the audio processing task by dividing it into distinct pathways: one for music recognition and one for speech recognition. The integration layer acts as a router that separates incoming audio input based on its type and directs it to the appropriate specialized module. This segmentation ensures that computational resources are allocated only to the necessary processing pathway, maintaining versatility while reducing energy consumption by avoiding simultaneous full processing of both types.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs partial processing by first analyzing audio input to determine its type before committing to full recognition processing. The integration layer performs a preliminary classification step that requires minimal computational resources compared to full speech or music recognition. This partial action approach ensures comprehensive service coverage is available while avoiding the excessive computational cost of running both recognition systems simultaneously on all audio input.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9620105B2Analyzing audio input for efficient speech and music recognition
Publication Date: 2017.04.11 APPLE INC
  • US9620105B2 patent drawing
  • US9620105B2 patent drawing
  • US9620105B2 patent drawing

AI summary

Systems and processes for analyzing audio input for efficient speech and music recognition are provided. In one example process, an audio input can be received. A determination can be made as to whether the audio input includes music. In addition, a determination can be made as to whether the audio input includes speech. In response to determining that the audio input includes music, an acoustic fingerprint representing a portion of the audio input that includes music is generated. In response to determining that the audio input includes speech rather than music, an end-point of a speech utterance of the audio input is identified.