Multi-Voice Speech Recognition Command Authentication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems face challenges in handling multiple voices in commands, failing to distinguish between parties and determine identity, leading to unsecure and limited functionality, resulting in poor user experiences and reduced usability.

Innovation Solution

The implementation of multi-voice speech recognition commands that detect and authenticate complex commands by generating hashes based on initial and subsequent triggers, allowing devices to identify entities and determine responses based on conversational context, enabling secure and adaptive interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional speech recognition systems are used, then the system structure is simple, but the system cannot distinguish between multiple voices and determine entity identity, leading to security issues

Engineering Contradiction:
ImprovesecurityVSAvoidsystem structure
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent segments the speech recognition process into distinct components: voice detection module, identity determination module, and command processing module. Each module handles a specific aspect of multi-voice recognition, allowing the system to differentiate between multiple speakers while maintaining manageable system complexity through functional decomposition.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary entity (such as a server or processing unit) that mediates between the multiple voice inputs and the command execution system. This intermediary analyzes voice characteristics, determines speaker identity, and authenticates commands, thereby enhancing security without requiring direct complex interactions between all system components.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If multi-voice speech recognition is implemented, then the adaptability and usability are enhanced, but the complexity of detecting and measuring multiple triggers increases

Engineering Contradiction:
Improvemulti-voice command handlingVSAvoidtrigger detection
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent applies preliminary action by pre-configuring voice profiles and authentication criteria before multi-voice recognition is needed. The system pre-processes voice samples to extract characteristic features and stores them for rapid comparison during actual command detection, reducing the real-time complexity of distinguishing between multiple speakers.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces complex mechanical or manual voice differentiation methods with automated signal processing and pattern recognition algorithms. By using computational techniques to analyze voice characteristics, extract features, and match patterns, the system reduces the difficulty of detecting multiple triggers while enhancing adaptability to different speakers and commands.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS12189744B2Techniques for multi-voice speech recognition commands
Publication Date: 2025.01.07 CAPITAL ONE SERVICES LLC
  • US12189744B2 patent drawing
  • US12189744B2 patent drawing
  • US12189744B2 patent drawing

AI summary

Various embodiments are generally directed to techniques for multi-voice speech recognition commands, such as based on monitoring a telecommunications channel between first and second devices, for instance. Some embodiments are particularly directed to prompting initiation of a transaction between a first entity associated with a first device and a second entity associated with a second device based on detection of an audible request corresponding to the second entity and an audible response corresponding to the first entity.