Joint Speaker and Content Model for Speech Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems are inefficient in real-time speaker verification and content recognition, as they require separate analyses and are often language-dependent, leading to high processing demands and limited applications.

Innovation Solution

A joint speaker and content model is used to analyze speech samples, allowing for simultaneous speaker verification and content recognition using a single input, leveraging phonetic and acoustic features to identify authorized users and commands without language specificity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If separate analyses are applied for speaker recognition and content recognition, then each analysis can focus on specific aspects of the speech signal, but the processing time and computational resources increase significantly

Engineering Contradiction:
Improverecognition accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent merges separate speaker recognition and content recognition analyses into a unified joint analysis framework. The system processes speech signals through a single integrated model that simultaneously extracts speaker-specific features and phonetic/lexical content features, eliminating the need for sequential processing and reducing overall computation time while maintaining recognition accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent develops a universal speech analysis model that performs multiple functions simultaneously - speaker verification, content recognition, and phonetic analysis - all within a single processing framework. This multi-functional approach allows the system to handle diverse recognition tasks without requiring separate specialized models for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If separate analyses are applied for speaker recognition and content recognition, then each model can specialize in different aspects, but the system complexity and processing resource requirements increase

Engineering Contradiction:
Improverecognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines multiple specialized models into a single integrated architecture that handles both speaker verification and content recognition. By merging the speaker recognition module and content analysis module into one unified system, the patent reduces the number of separate processing pipelines and simplifies the overall system structure while preserving the specialized capabilities needed for accurate recognition.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent segments the speech signal processing into distinct feature extraction pathways within the unified model - one pathway for speaker-specific characteristics and another for phonetic/lexical content - allowing each aspect to be analyzed independently within the integrated framework, thereby maintaining specialization without requiring separate external systems.

Inventive Principle:
Principle #1Segmentation

3Adaptability or versatility

If traditional neural networks are used for speech analysis, then the model size and depth are limited, but the system cannot capture complex phonetic and speaker-specific features simultaneously

Engineering Contradiction:
Improvefeature extraction capabilityVSAvoidmodel size
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent employs a deep neural network with dynamic architecture that can adaptively adjust its processing depth and complexity based on the specific recognition task. The network dynamically routes speech signals through different processing pathways - emphasizing speaker verification layers or content recognition layers - allowing the model to optimize its computational resources for the current task while maintaining the capability to handle complex feature extraction.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10476872B2Joint speaker authentication and key phrase identification
Publication Date: 2019.11.12 SRI INTERNATIONAL
  • US10476872B2 patent drawing
  • US10476872B2 patent drawing
  • US10476872B2 patent drawing

AI summary

A spoken command analyzer computing system includes technologies configured to analyze information extracted from a speech sample and, using a joint speaker and phonetic content model, both determine whether the analyzed speech includes certain content (e.g., a command) and to identify the identity of the human speaker of the speech. In response to determining that the identity matches the authorized user's identity and determining that the analyzed speech includes the modeled content (e.g., command), an action corresponding to the verified content (e.g., command) is performed by an associated device.