Map Pronunciation Annotation via User Speech Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing mapping applications and location-related services often struggle to accurately pronounce location names, especially in foreign countries, due to incomplete and inaccurate audio databases that lack robust solutions for various geographical locations.

Innovation Solution

A system that allows users to upload their own sound recordings of location names, generates a speech model based on these recordings, and selects the most typical pronunciation using Hidden Markov Models or other speech processing techniques to provide accurate audio output.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing audio databases are used to pronounce location names, then the system has simple database storage and retrieval, but the pronunciation accuracy and coverage are insufficient

Engineering Contradiction:
Improvepronunciation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system enables users to autonomously upload their own pronunciations of location names, and users can select their preferred pronunciations from multiple options. This self-service approach allows the system to accumulate diverse pronunciation data without requiring manual curation, thereby improving pronunciation accuracy and coverage while avoiding the complexity of manual database maintenance

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system pre-processes uploaded pronunciations by storing them in a database and making them available for selection before actual use. By preparing pronunciation data in advance through user uploads and organizing it in a searchable format, the system ensures accurate pronunciations are ready when needed, resolving the contradiction between accuracy and complexity

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If a fixed audio database is used, then the system structure is simple, but the adaptability to various geographical locations is limited

Engineering Contradiction:
Improvecoverage of geographical locationsVSAvoiddata collection and processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The pronunciation database transitions from a static fixed structure to a dynamic user-contributed system. Users can continuously add new pronunciations for different geographical locations, and the system adapts by storing and managing these variable inputs. This dynamic approach enables the system to expand coverage to diverse geographical locations while managing complexity through automated user-driven data collection

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system serves multiple geographical locations and pronunciation variants through a single unified platform. By accepting uploads from any user for any location and providing selection capabilities, the system achieves universal coverage across diverse geographical contexts without requiring separate systems for each location, thereby improving adaptability while controlling complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8949125B1Annotating maps with user-contributed pronunciations
Publication Date: 2015.02.03 GOOGLE LLC
  • US8949125B1 patent drawing
  • US8949125B1 patent drawing
  • US8949125B1 patent drawing

AI summary

Systems and methods are provided to select a most typical pronunciation of a location name on a map from a plurality of user pronunciations. A server generates a reference speech model based on user pronunciations, compares the user pronunciations with the speech model and selects a pronunciation based on comparison. Alternatively, the server compares the distance between one the user pronunciations and every other user pronunciations and selects a pronunciation based on comparison. The server then annotates the map with the selected pronunciation and provides the audio output of the location name to a user device upon a user's request.