Acoustic Model Generation Using Social Network Demographics

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems face limitations in accurately interpreting spoken words due to variations in pronunciation based on demographic characteristics such as age, gender, and regional accent, leading to suboptimal performance in distinguishing homonyms and device-dependent acoustic models.

Innovation Solution

The method involves generating customized acoustic models by collecting and combining speech data from users with similar demographic characteristics through social networking sites, using social graph information to identify and cluster users based on demographics like age, gender, and regional accent, and associating these models with target demographic characteristics.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional speech recognition systems use general acoustic models, then device complexity is reduced, but measurement precision deteriorates due to variations in pronunciation based on demographic characteristics

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidacoustic model complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the general acoustic model into multiple demographic-specific acoustic models tailored to different user groups (e.g., age groups, gender, regional accents). Each segmented model focuses on specific pronunciation patterns of its target demographic, improving recognition accuracy without requiring the system to maintain one monolithic complex model.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by creating acoustic models with specialized characteristics for specific demographic groups. Each acoustic model is optimized with local pronunciation features relevant to its target demographic (e.g., Southern accent patterns for specific regions, age-specific intonation patterns), rather than using a uniform model for all users.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If speech recognition systems collect and process more speech data from diverse users, then measurement precision improves, but loss of information increases due to difficulty in managing and associating data with correct demographic characteristics

Engineering Contradiction:
Improveacoustic model accuracyVSAvoiddemographic characteristic association accuracy
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent introduces social networking data as an intermediary mechanism to bridge speech data and demographic characteristics. Instead of directly collecting and matching demographic information with speech samples, the system uses social graph relationships and profile information from social networks to infer and associate demographic characteristics with speech data, reducing information loss.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements feedback loops where the system continuously refines demographic characteristic associations based on recognition performance. When speech recognition accuracy is measured, this feedback is used to improve the association between speech data and demographic characteristics, creating a self-correcting system that reduces information loss over time.

Inventive Principle:
Principle #23Feedback

3Adaptability or versatility

If the system creates customized acoustic models for multiple demographic characteristics, then adaptability improves, but device complexity increases due to managing multiple models and social graph data

Engineering Contradiction:
Improvedemographic customization capabilityVSAvoidsystem architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent makes the acoustic model system universal by designing it to handle multiple demographic characteristics (age, gender, region, accent) through a unified framework. The same infrastructure processes social graph data, generates demographic profiles, and manages acoustic models for various user groups, allowing the system to adapt to different demographics without requiring separate independent systems for each characteristic.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent applies nested doll by organizing acoustic models hierarchically within the social graph structure. Demographic-specific acoustic models are nested within broader demographic categories, which are themselves nested within the overall social network framework. This nested organization allows the system to manage multiple levels of customization (from broad regional accents to specific age groups) in a structured manner.

Inventive Principle:
Principle #7Nested doll (Nesting)

Data Source

PatentUS9460716B1Using social networks to improve acoustic models
Publication Date: 2016.10.04 GOOGLE LLC
  • US9460716B1 patent drawing
  • US9460716B1 patent drawing
  • US9460716B1 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for acoustic model generation. One of the methods includes identifying one or more demographic characteristics for a user of a social networking site. The method includes receiving speech data from the user, the speech data associated with a user device. The method includes storing the speech data associated with demographic characteristics of the user and the user device.