Speech Processing Apparatus with Pre-constructed Response Database

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition services incur high usage costs and network delays due to the need for extensive processing and data transmission for every query, leading to inefficient use of hardware resources and increased network data usage.

Innovation Solution

A speech processing method and apparatus that preconstructs frequently used query and response utterances in a database, allowing for direct retrieval and updating based on usage frequency, storage capacity, and predetermined thresholds, thereby reducing the need for real-time processing and data transmission.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech recognition service processes every query through complete speech-to-text conversion, natural language understanding, and text-to-speech synthesis, then comprehensive service coverage is achieved, but usage costs and network data consumption increase significantly

Engineering Contradiction:
Improveservice coverageVSAvoidusage cost
Core Design Contradiction:
Adaptability or versatilityVSLoss of energy

Solution Approach 1:

The patent pre-constructs spoken response utterances in the database corresponding to frequently queried topics before actual use. When a user asks a question about a pre-constructed topic, the system directly retrieves and plays the pre-synthesized response without performing real-time speech-to-text conversion, natural language understanding, or text-to-speech synthesis, thereby avoiding repeated processing costs and network data consumption while maintaining comprehensive service coverage

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies different processing strategies to different types of queries: for frequently queried topics with pre-constructed responses, it uses direct database retrieval and playback; for other queries, it performs complete speech recognition processing. This localized quality adjustment optimizes resource allocation by applying intensive processing only where necessary while using efficient retrieval for common queries

Inventive Principle:
Principle #3Local quality

2Reliability

If speech recognition service performs complete processing for every query, then accurate responses are generated, but network delay increases

Engineering Contradiction:
Improveresponse accuracyVSAvoidnetwork delay
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system pre-generates and stores spoken response utterances in the database before actual queries occur. When a user asks a question about a pre-constructed topic, the response is immediately retrieved and played back without undergoing real-time speech-to-text conversion, natural language understanding, or text-to-speech synthesis, thereby eliminating processing delays while ensuring accurate responses through pre-computed results

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If all query-text to spoken response utterance sets are retained in the database, then complete coverage is maintained, but hardware resources are wasted

Engineering Contradiction:
Improvecoverage completenessVSAvoidstorage resource
Core Design Contradiction:
Adaptability or versatilityVSQuantity of substance

Solution Approach 1:

The patent implements a frequency-based retention strategy where query-text to spoken response utterance sets are retained in the database only if their providing frequency exceeds a predetermined threshold. Sets with low providing frequencies are deleted to free up storage resources. This approach maintains complete coverage for frequently queried topics while efficiently discarding rarely used data, thereby optimizing the balance between coverage completeness and hardware resource utilization

Inventive Principle:
Principle #34Discarding and recovering

Data Source

PatentUS11328718B2Speech processing method and apparatus therefor
Publication Date: 2022.05.10 LG ELECTRONICS INC
  • US11328718B2 patent drawing
  • US11328718B2 patent drawing
  • US11328718B2 patent drawing

AI summary

A speech processing method and a speech processing apparatus which execute a mounted artificial intelligence (AI) algorithm and/or machine learning algorithm to perform speech processing so that electronic devices and a server may communicate with each other in a 5G communication environment are disclosed. A speech processing method according to an exemplary embodiment of the present disclosure may include collecting a user's spoken utterance including a query, generating a query text as a text conversion result for the user's spoken utterance including a query, searching whether there is a query text-spoken response utterance set including a spoken response utterance for the query text in a database which is constructed in advance, and when there is a query text-spoken response utterance set including a spoken response utterance for the query text in the database, providing the spoken response utterance included in the query text-spoken response utterance set.