Speech Analysis Apparatus for Conference Keyword Visualization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies for analyzing speech in groups only indicate the degree of contribution based on speaking time and do not display the actual content of speech during debates or conferences.
Innovation Solution
An information processing apparatus and method that includes a detector for identifying utterances, a textualization device for converting speech to text, a keyword detector, and a display controller to show detected keywords, allowing for the visualization of speech content in a group setting.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If only speaking time duration is used to measure contribution, then the measurement is simple, but the content of speech is not captured
Solution Approach 1:
The speech analysis system segments the continuous audio data into discrete utterances, then further segments each utterance into individual words. This segmentation allows the system to analyze speech content at multiple levels (utterance level for speaker identification, word level for keyword extraction) without requiring complex simultaneous processing of the entire speech stream.
Solution Approach 2:
The patent introduces text data as an intermediary between audio data and keyword detection. The speech recognition device converts audio utterances into text, which then serves as the input for keyword detection. This intermediary text representation simplifies the keyword detection process compared to direct audio analysis, while preserving the semantic content needed for meaningful analysis.
2Loss of information
If speech content is analyzed in real-time, then the information is useful, but the processing time increases
Solution Approach 1:
The system performs preliminary actions by first detecting utterances and converting them to text before keyword detection. This preliminary textualization prepares the data in advance for efficient keyword searching, allowing the keyword detection to operate on already-processed text rather than raw audio, reducing the time penalty for content analysis.
Solution Approach 2:
The patent applies partial action by detecting only predetermined keywords rather than analyzing every word in the speech. This selective approach focuses computational resources on extracting meaningful keywords that are relevant to the analysis goal, rather than processing the entire speech content, thus reducing processing time while maintaining information usefulness.
3Measurement precision
If all utterances are converted to text, then complete content analysis is possible, but the processing load increases
Solution Approach 1:
The system extracts only the essential elements for analysis: predetermined keywords that are relevant to the debate or conference topic. Rather than converting and analyzing all speech content, the extraction of specific keywords provides sufficient measurement precision for evaluating speech contributions while significantly reducing the processing load compared to comprehensive text conversion of all utterances.
Solution Approach 2:
The patent changes the parameter of analysis from complete utterance text to specific keyword presence/absence. This parameter change from analyzing all text content to detecting specific keywords maintains the ability to measure speech contribution precision (by tracking relevant topic keywords) while improving processing efficiency through reduced computational requirements.
Data Source
AI summary
An information processing apparatus includes a first detector, a textualization device, a second detector, a display device and a display controller. The first detector detects, from audio data in which speech of each person in a group composed of a plurality of persons has been recorded, each utterance made during the speech. The textualization device converts contents of each utterance detected by the first detector into text. The second detector detects predetermined keywords included in each utterance on the basis of text data obtained through textualization by the textualization device. The display controller causes the display device to display the predetermined keywords detected by the second detector.


