Speech Processing Apparatus with Pre-constructed Response Database
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech recognition services incur high usage costs and network delays due to the need for extensive processing and data transmission for every query, leading to inefficient use of hardware resources and increased network data usage.
Innovation Solution
A speech processing method and apparatus that preconstructs frequently used query and response utterances in a database, allowing for direct retrieval and updating based on usage frequency, storage capacity, and predetermined thresholds, thereby reducing the need for real-time processing and data transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech recognition service processes every query through complete speech-to-text conversion, natural language understanding, and text-to-speech synthesis, then comprehensive service coverage is achieved, but usage costs and network data consumption increase significantly
Solution Approach 1:
The patent pre-constructs spoken response utterances in the database corresponding to frequently queried topics before actual use. When a user asks a question about a pre-constructed topic, the system directly retrieves and plays the pre-synthesized response without performing real-time speech-to-text conversion, natural language understanding, or text-to-speech synthesis, thereby avoiding repeated processing costs and network data consumption while maintaining comprehensive service coverage
Solution Approach 2:
The patent applies different processing strategies to different types of queries: for frequently queried topics with pre-constructed responses, it uses direct database retrieval and playback; for other queries, it performs complete speech recognition processing. This localized quality adjustment optimizes resource allocation by applying intensive processing only where necessary while using efficient retrieval for common queries
2Reliability
If speech recognition service performs complete processing for every query, then accurate responses are generated, but network delay increases
Solution Approach 1:
The system pre-generates and stores spoken response utterances in the database before actual queries occur. When a user asks a question about a pre-constructed topic, the response is immediately retrieved and played back without undergoing real-time speech-to-text conversion, natural language understanding, or text-to-speech synthesis, thereby eliminating processing delays while ensuring accurate responses through pre-computed results
3Adaptability or versatility
If all query-text to spoken response utterance sets are retained in the database, then complete coverage is maintained, but hardware resources are wasted
Solution Approach 1:
The patent implements a frequency-based retention strategy where query-text to spoken response utterance sets are retained in the database only if their providing frequency exceeds a predetermined threshold. Sets with low providing frequencies are deleted to free up storage resources. This approach maintains complete coverage for frequently queried topics while efficiently discarding rarely used data, thereby optimizing the balance between coverage completeness and hardware resource utilization
Data Source
AI summary
A speech processing method and a speech processing apparatus which execute a mounted artificial intelligence (AI) algorithm and/or machine learning algorithm to perform speech processing so that electronic devices and a server may communicate with each other in a 5G communication environment are disclosed. A speech processing method according to an exemplary embodiment of the present disclosure may include collecting a user's spoken utterance including a query, generating a query text as a text conversion result for the user's spoken utterance including a query, searching whether there is a query text-spoken response utterance set including a spoken response utterance for the query text in a database which is constructed in advance, and when there is a query text-spoken response utterance set including a spoken response utterance for the query text in the database, providing the spoken response utterance included in the query text-spoken response utterance set.


