Personalized Speech Recognition via Local Communication Networks

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems for mobile information search suffer from low accuracy due to open vocabulary perplexity, especially in mobile information access, where users face difficulties typing on small screens and keyboards, limiting the practicality of speech input for information retrieval.

Innovation Solution

The system employs a local communication network to generate personalized and adaptive speech recognition models by analyzing a user's calling history, identifying local neighborhoods, and creating language models based on frequently called entities, which reduces word perplexity and improves search accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speech recognition systems use open vocabulary for mobile information search, then users can perform searches using speech input without typing, but the recognition accuracy deteriorates due to word perplexity

Engineering Contradiction:
Improveease of speech inputVSAvoidspeech recognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent applies local quality by creating personalized language models for each user based on their local communication network (calling history, contacts, text messages). Instead of using a generic open vocabulary model, the system adapts the speech recognition to each user's specific linguistic patterns, frequently contacted entities, and communication habits, thereby improving recognition accuracy while maintaining speech input convenience

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system performs preliminary action by pre-processing and analyzing user communication data (calling history, contacts, messages) to build personalized language models before speech recognition is needed. This pre-computation of user-specific language patterns enables more accurate real-time speech recognition without requiring typing

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If speech recognition models are personalized for each user, then recognition accuracy improves, but system complexity increases due to model generation and management

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system applies self-service by automatically generating personalized language models using the user's own communication data (calling history, contacts, messages) without requiring manual configuration or user intervention. The system self-adapts to each user's linguistic patterns and frequently contacted entities, reducing the perceived complexity for users while maintaining high recognition accuracy

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent uses a universal framework that handles multiple functions: it processes calling history, contacts, and text messages to build language models, and can serve multiple users with the same system infrastructure. This multi-functional approach manages complexity by reusing common processing pipelines across different users while still providing personalized recognition

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS7966171B2System and method for increasing accuracy of searches based on communities of interest
Publication Date: 2011.06.21 AT&T INTELLECTUAL PROPERTY II LP
  • US7966171B2 patent drawing
  • US7966171B2 patent drawing
  • US7966171B2 patent drawing

AI summary

Disclosed are systems, methods and computer-readable media for using a local communication network to generate a speech model. The method includes retrieving for an individual a list of numbers in a calling history, identifying a local neighborhood associated with each number in the calling history, truncating the local neighborhood associated with each number based on the at least one parameter, retrieving a local communication network associated with each number in the calling history and each phone number in the local neighborhood, and creating a language model for the individual based on the retrieved local communication network. The generated language model may be used for improved automatic speech recognition for audible searches as well as other modules in a spoken dialog system.