Speaker Identification System Using Acoustic Prosodic Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional CAPTCHA systems, including visual and audio-based ones, are vulnerable to sophisticated machine vision and speech recognition technologies, making them ineffective in distinguishing between human and machine inputs, particularly as machines improve in recognizing and synthesizing human-like inputs.

Innovation Solution

A method that involves receiving speech utterances related to randomly selected challenge text, processing them to compute acoustical characteristics, and determining whether the utterance originated from a human or a machine by comparing these characteristics with reference sets for human and computer synthesized voices, while also considering prosodic elements and using a combination of visual and audio challenges to enhance differentiation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional visual CAPTCHA systems are used, then they are easy to implement and understand, but they become vulnerable to sophisticated machine vision systems that can break them

Engineering Contradiction:
Improveeffectiveness in distinguishing human from machineVSAvoidcomplexity of differentiation system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system segments the differentiation task into multiple independent analysis channels: acoustical characteristic analysis, prosodic characteristic analysis, and challenge-response verification. Each channel processes specific features separately and their results are combined to make the final determination, improving reliability without requiring a single complex system

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from analyzing only visual dimensions to incorporating acoustic and prosodic dimensions. By adding these new dimensions of analysis (acoustical characteristics, prosodic characteristics), the system creates a multi-dimensional verification space that is much harder for machines to spoof while maintaining implementation feasibility

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If audio CAPTCHA systems are used, then they provide an alternative to visual challenges, but speech recognizers are improving rapidly making them vulnerable to machine breakdown

Engineering Contradiction:
Improvesecurity against machine recognitionVSAvoidadaptability to machine speech recognition improvements
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The system performs preliminary analysis of acoustical and prosodic characteristics before final verification. By extracting and analyzing these fundamental characteristics early in the process, the system establishes baseline expectations for human speech patterns that can adapt to evolving machine speech recognition capabilities

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system monitors and adapts to changes in speech recognition parameters by continuously analyzing acoustical and prosodic features. When machine recognition capabilities evolve, the system can adjust its parameter thresholds and analysis criteria based on the observed changes in speech patterns, maintaining security against new threats

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If speech utterances are analyzed for acoustical characteristics, then differentiation between human and machine voices improves, but the processing complexity increases

Engineering Contradiction:
Improveprecision in voice origin determinationVSAvoidcomplexity of speech processing system
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system extracts only the most critical acoustical and prosodic characteristics from speech utterances rather than analyzing the entire speech signal. By selecting and extracting specific key features (such as fundamental frequency, formant frequencies, jitter, shimmer), the system achieves high measurement precision while keeping processing complexity manageable

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system performs partial analysis by focusing on specific acoustical and prosodic parameters rather than comprehensive speech analysis. This selective approach analyzes only the most discriminative features needed for human-machine differentiation, achieving sufficient precision without the complexity of full speech processing

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10013972B2System and method for identifying speakers
Publication Date: 2018.07.03 KNAPP INVESTMENT COMPANY LIMITED
  • US10013972B2 patent drawing
  • US10013972B2 patent drawing
  • US10013972B2 patent drawing

AI summary

An electronic challenge system is used to control access to resources by using a spoken test to identify an origin of a voice. The test is based on a series of questions posed during an interactive dialog session with the entity attempting access.