Keyword Extraction from Speech-to-Text Call Logs
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional call logs on mobile devices only provide basic information such as contacts and timestamps, lacking context about the conversations, which can make it difficult for users to recall the details of past communications.
Innovation Solution
A computing device performs speech-to-text processing and keyword extraction on audio communications, generating keywords that are relevant to the conversation and associates them with contact information, allowing for a graphical display of these keywords in the communication log, enabling users to quickly recall conversation topics.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If traditional call logs only store basic information (contacts and timestamps), then the device complexity is low and storage requirements are minimal, but the information completeness and user ability to recall conversation details is insufficient
Solution Approach 1:
The system extracts only the most relevant keywords from speech inputs rather than storing complete transcriptions. Keyword extraction processing identifies and stores only salient terms that represent conversation topics, significantly reducing information storage requirements while maintaining the ability to recall conversation context.
Solution Approach 2:
The conversation information is segmented into discrete keywords rather than storing continuous text. This segmentation allows the system to store manageable units of information (individual keywords) that can be efficiently processed, stored, and retrieved without requiring complex storage infrastructure.
2Loss of information
If speech-to-text processing and keyword extraction are performed on all audio communications, then the information completeness and recall ability improve, but the processing time and computational resources increase
Solution Approach 1:
The system performs keyword extraction rather than complete speech-to-text transcription. By extracting only key terms and phrases that represent conversation essence, the processing time is significantly reduced while still providing sufficient context for users to recall conversation details.
Solution Approach 2:
The system performs partial processing by extracting only the most important keywords rather than transcribing the entire conversation. This partial action approach balances processing time constraints with the need to retain sufficient conversation context for effective recall.
3Ease of operation
If keywords are stored and displayed in association with contact information, then the ease of operation and user experience improve, but the device complexity and data management requirements increase
Solution Approach 1:
The system merges keywords with existing contact information in the communication log. By combining conversation keywords with contact data that users already have organized, the system enhances recall ability without requiring separate complex data management structures for keywords.
Solution Approach 2:
The keyword storage system leverages the existing communication log infrastructure. Rather than creating a separate complex data management system, the keywords are integrated into the universal contact and communication record structure that already exists in mobile devices.
Data Source
AI summary
A computing device may extract keywords from a phone call or other audio communication and later display those keywords in a call log or in a caller ID. In one example, a method performed by at least one processor of a first computing device includes receiving speech inputs during an audio communication between the first computing device and a second device. The method further includes performing speech-to-text processing on the speech inputs to generate text based on the speech inputs. The method further includes performing keyword extraction processing on the text to generate one or more keywords based on the text, that score highly as relevant indicators of the audio communication, based on one or more keyword extraction criteria. The method further includes storing the one or more keywords in association with identifying information associated with the second device. The method further includes outputting the one or more keywords, in association with the identifying information associated with the second device, at a display of the first computing device.


