Voice-Activated Product Audio Retrieval via Biased Speech Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Users, especially children and individuals with speech disabilities, face challenges in interacting with voice-activated products due to limited availability of audible content and difficulties in clear speech recognition, and there is a need for a system to record and retrieve audio narratives effectively.

Innovation Solution

A voice-activated product that allows users to record and store audio content with associated identifiers, enabling biased speech recognition and easy retrieval, even for users with unclear speech, by incorporating bibliographic information and sound effects, and providing interactive dialogs tailored for different users.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If standard speech recognition is used for voice-activated products, then the system can process voice commands, but users with unclear speech (children, individuals with disabilities) cannot effectively interact

Engineering Contradiction:
Improveease of voice interactionVSAvoidspeech recognition accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system changes the speech recognition parameters by biasing towards specific identifiers (titles, names) based on audio characteristics analysis. When detecting that a user has unclear speech patterns, the system adjusts recognition sensitivity and priority to favor known identifiers, thereby improving reliability for vulnerable users while maintaining ease of operation.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system implements feedback by analyzing audio characteristics of voice inputs and using this information to adjust speech recognition behavior dynamically. The feedback loop detects speech clarity levels and modifies recognition strategies accordingly, creating an adaptive system that improves accuracy for users with unclear speech without requiring them to change their speaking style.

Inventive Principle:
Principle #23Feedback

2Quantity of substance

If audio content is made available on voice-activated products, then users have access to narratives, but availability is limited and requires purchases

Engineering Contradiction:
Improvequantity of audio contentVSAvoidcontent acquisition cost
Core Design Contradiction:
Quantity of substanceVSEase of manufacture

Solution Approach 1:

The system enables self-service content creation by allowing users to record their own audio narratives using their personal devices. Users can store these recordings locally or in the cloud without requiring platform purchases. The voice-activated product simply provides the interface to access and play back user-generated content, eliminating the need for content acquisition costs while significantly increasing available content quantity.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system makes the voice-activated product universal by enabling it to serve multiple functions: playing pre-loaded content, accessing user-recorded content, and facilitating content creation. This multi-functionality increases the quantity of available audio content without requiring separate purchases for different content types, as the same platform handles all content delivery methods.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of operation

If users record and store audio content with identifiers, then retrieval becomes easier, but the system complexity increases

Engineering Contradiction:
Improvecontent retrieval easeVSAvoidsystem complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system applies preliminary action by having users provide identifiers (titles, names, descriptions) when recording audio content. These identifiers are stored and indexed in advance, creating a searchable database. During playback, users can simply query by these pre-provided identifiers without needing to navigate complex interfaces, thereby improving retrieval ease while the system handles the complexity of storage and indexing automatically.

Inventive Principle:
Principle #10Preliminary action

4Adaptability or versatility

If voice-activated products provide tailored interactive dialogs, then accessibility is enhanced for different users, but the device complexity increases

Engineering Contradiction:
Improveuser-specific adaptabilityVSAvoiddialog management complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system changes interaction parameters dynamically based on user characteristics. By analyzing audio characteristics and identifying user types (children, adults, individuals with disabilities), the system adjusts dialog complexity, vocabulary, and interaction patterns accordingly. This parameter adaptation enables tailored interactive dialogs that enhance accessibility for different user groups while managing complexity through automated detection and adjustment rather than manual configuration.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11238854B2Facilitating creation and playback of user-recorded audio
Publication Date: 2022.02.01 GOOGLE LLC
  • US11238854B2 patent drawing
  • US11238854B2 patent drawing
  • US11238854B2 patent drawing

AI summary

Methods, apparatus, and computer readable media are described related to recording, organizing, and making audio files available for consumption by voice-activated products. In various implementations, in response to receiving an input from a first user indicating that the first user intends to record audio content, audio content may be captured and stored. Input may be received from the first user indicating at least one identifier for the audio content. The stored audio content may be associated with the at least one identifier. A voice input may be received from a subsequent user. In response to determining that the voice input has particular characteristics, speech recognition may be biased in respect of the voice input towards recognition of the at least one identifier. In response to recognizing, based on the biased speech recognition, presence of the at least one identifier in the voice input, the stored audio content may be played.