Offline Speech Output Optimization Using Local Database Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing offline text to speech solutions for vehicle-mounted and mobile terminals result in mechanical-sounding speech due to limited processing performance and storage, impacting user experience, especially during navigation.

Innovation Solution

A method that uses a local speech database for text matching to determine high-quality human-like speech output, splitting target text into keywords for matching and splicing speech segments to optimize output speech quality without relying solely on offline text to speech synthesis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If offline text to speech is used due to limited processing performance and storage space, then the device can operate independently without network connection, but the speech output sounds mechanical and lacks naturalness

Engineering Contradiction:
Improveoffline operation capabilityVSAvoidspeech naturalness
Core Design Contradiction:
Adaptability or versatilityVSEase of manufacture

Solution Approach 1:

The patent segments the speech output process into two parts: using a pre-configured local speech database for common navigation instructions to provide natural speech, and using offline text-to-speech synthesis only when necessary for uncommon phrases. This segmentation resolves the contradiction by prioritizing naturalness for frequent use cases while maintaining offline independence.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies preliminary action by pre-configuring a local speech database with high-quality human speech recordings for common navigation instructions before offline operation begins. This allows the device to immediately provide natural speech output for common phrases without requiring network connection or complex real-time processing.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If a large storage space is allocated for offline text to speech programs, then more comprehensive speech functions are available, but the device's limited storage space cannot accommodate the program package

Engineering Contradiction:
Improvespeech function comprehensivenessVSAvoidstorage space consumption
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent applies local quality by creating a specialized local speech database that contains only high-quality human speech recordings for common navigation instructions, rather than storing complete offline text-to-speech programs. This selective approach provides comprehensive natural speech for frequent use cases while consuming minimal storage space.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent uses copying by storing pre-recorded human speech samples in the local database rather than storing the complete text-to-speech synthesis system. This allows the device to access natural speech output for common phrases without requiring the full synthesis program package, thus reducing storage requirements.

Inventive Principle:
Principle #26Copying

3Reliability

If offline text to speech is used to ensure speech output in offline state, then the device can function without network connection, but the speech output lacks anthropomorphic characteristics

Engineering Contradiction:
Improvespeech output availabilityVSAvoidanthropomorphic quality
Core Design Contradiction:
ReliabilityVSEase of manufacture

Solution Approach 1:

The patent segments speech output into two sources: the local speech database for common instructions providing anthropomorphic quality, and offline text-to-speech for uncommon phrases ensuring availability. This segmentation maintains reliability for all offline scenarios while prioritizing anthropomorphic quality for the most common use cases.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces the local speech database as an intermediary between the user's offline operation needs and the requirement for natural speech output. This intermediary provides pre-configured high-quality speech recordings that bridge the gap between availability and anthropomorphic quality without requiring the full text-to-speech system.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP3882909B1Speech output method and apparatus, device and medium
Publication Date: 2024.02.21 APOLLO INTELLIGENT CONNECTIVITY (BEIJING) TECH CO LTD
  • EP3882909B1 patent drawingFigure 1
  • EP3882909B1 patent drawingFigure 2
  • EP3882909B1 patent drawingFigure 3

AI summary

Embodiments of the present disclosure disclose a speech output method and apparatus. The method includes: determining (S101) a target text to be processed; matching (S102) the target text with a local text database to determine a preset text corresponding to the target text; and determining (S103), based on the preset text, output speech of the target text from a local speech database to output the output speech; wherein the local speech database is pre-configured based on a correspondence between a text and speech. According to embodiments of the present disclosure, the output speech may be optimized when a device supporting speech interaction is offline so as to improve an anthropomorphic level of the output speech and to mitigate the impact of mechanical speech on user experience.