ASR Audio Quality Metrics for Mobile Subsystem Testing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Audio distortion caused by clipping, lost samples, or microphone frequency response non-linearity in mobile devices can significantly impact speech recognition accuracy, and existing solutions face challenges due to communication problems and corporate boundaries.

Innovation Solution

A method that allows manufacturers to enhance mobile device audio subsystems by recording standardized audio inputs, sending them to an automated speech recognition engine for processing, and generating audio quality metrics to assist in reconfiguration or redesign, thereby alleviating the burden on ASR or search engine operators.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manufacturers collaborate on audio subsystem design, then audio quality can be improved, but communication problems and corporate boundaries prevent effective collaboration

Engineering Contradiction:
Improveaudio qualityVSAvoidcollaboration complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces an automated speech recognition engine as an intermediary that objectively evaluates audio subsystem performance. This mediator processes audio recordings from mobile devices and generates standardized quality metrics, enabling manufacturers to improve audio quality without direct communication or collaboration, thus resolving the contradiction between improving audio quality and avoiding collaboration complexity.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If ASR engine operators test each new mobile device for compatibility, then speech recognition accuracy can be ensured, but this creates a significant operational burden

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidtesting efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent implements preliminary action by having mobile device manufacturers perform audio subsystem testing and send recordings to the ASR engine before devices are widely deployed. This allows the ASR engine to pre-evaluate audio quality and generate feedback metrics, ensuring speech recognition accuracy is assessed in advance without requiring operators to test each new device individually, thus resolving the contradiction between measurement precision and productivity.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If manufacturers test and enhance audio subsystems independently, then testing burden on ASR operators is reduced, but audio quality issues may persist without proper feedback

Engineering Contradiction:
Improvetesting efficiencyVSAvoidaudio quality feedback
Core Design Contradiction:
ProductivityVSLoss of information

Solution Approach 1:

The patent implements a feedback mechanism where the ASR engine analyzes audio recordings from mobile devices and generates standardized audio quality metrics that are relayed back to manufacturers. This feedback loop enables manufacturers to independently test and enhance audio subsystems while receiving actionable quality information, resolving the contradiction between testing efficiency and preventing loss of audio quality feedback.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8983845B1Third-party audio subsystem enhancement
Publication Date: 2015.03.17 GOOGLE LLC
  • US8983845B1 patent drawing
  • US8983845B1 patent drawing
  • US8983845B1 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for performing audio subsystem enhancement. In one aspect, a method includes: receiving a voice search query by an automatic speech recognition (ASR) engine that processes voice search queries for a search engine, wherein the voice search query includes an audio signal that corresponds to an utterance, and a test flag that indicates that an audio test is being performed; performing speech recognition on the audio signal to select one or more textual, candidate transcriptions that match the utterance; generating, in response to receiving the test flag, one or more audio quality metrics using the audio signal; and generating a response to the voice search query by the ASR engine, wherein the response references one or more of the candidate transcriptions and one or more of the audio quality metrics.