Audio-Activated Resource Access via Speaker Voiceprint Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for accessing resources are limited by the need for visual encoding in optical patterns, which is impractical for non-visual communication media like audio, and cannot seamlessly direct users to relevant content or applications without manual intervention.

Innovation Solution

The implementation of audio-activated resource access using speaker recognition systems, where a user device captures audio, identifies speakers, and transmits identifiers to a server to retrieve corresponding resources from a database, enabling seamless navigation to relevant content or applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If optical encoding methods (barcodes, QR codes) are used for resource access, then visual information can be captured and processed, but these methods are not applicable to non-visual communication media and require manual scanning intervention

Engineering Contradiction:
Improveapplicability to different communication mediaVSAvoidmanual intervention requirement
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent replaces manual mechanical scanning operations with automated audio-based resource access. The system automatically captures audio, processes it through speaker recognition, and retrieves resources without requiring manual scanning or visual interaction, thus substituting the mechanical scanning process with an automated audio-driven system.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Solution Approach 2:

The patent creates a universal resource access system that works across multiple communication media types. By using audio as the input modality that can be extracted from various sources (television, radio, audio streams, live speech), the system achieves versatility across visual and non-visual media, replacing media-specific access methods with a unified audio-based approach.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Extent of automation

If speaker recognition systems are implemented for automated resource access, then manual intervention is eliminated and adaptability to audio media is achieved, but system complexity increases

Engineering Contradiction:
Improveautomated resource accessVSAvoidsystem architecture complexity
Core Design Contradiction:
Extent of automationVSDevice complexity

Solution Approach 1:

The patent introduces a server system as an intermediary that handles the complex speaker recognition and resource matching operations. Instead of embedding complex speaker recognition algorithms directly in user devices, the system offloads this complexity to a centralized server that processes audio captures and returns resource identifiers, thus reducing device complexity while maintaining high automation.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent divides the resource access system into separate functional modules: audio capture at the user device, speaker recognition and resource matching at the server, and resource delivery back to the user device. This segmentation allows each component to be optimized independently and reduces the complexity burden on individual devices by distributing system functions across multiple components.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS9996628B2Providing audio-activated resource access for user devices based on speaker voiceprint
Publication Date: 2018.06.12 VERISIGN INC
  • US9996628B2 patent drawing
  • US9996628B2 patent drawing
  • US9996628B2 patent drawing

AI summary

This disclosure includes, for example, methods and computer systems for providing audio-activated resource access for user devices. The computer systems may store instructions to cause the processor to perform operations, comprising capturing audio at a user device. The operations may also comprise using a speaker recognition system to identify a speaker in the transmitted audio and/or using a speech-to-text converter to identify text in the captured audio. The speaker identity or a condensed version of the speaker identity or other metadata along with the speaker identity may be transmitted to a server system to determine a corresponding speaker identity entry. The operations may also comprise receiving a resource corresponding to the identified speaker entry in the server system.