Audio Signal Processing for User Intent Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing information processing systems struggle to read and interpret user emotions and attitudes beyond the content of speech, as they primarily focus on linguistic information and lack the ability to analyze paralinguistic cues.
Innovation Solution
An information processing apparatus and method that includes an audio signal acquisition section, a time identification section, and an output section, which identifies 'evaluation target time' by distinguishing between soundless and filler times in user audio input, allowing for the consideration of thinking time and emotional cues in determining agent responses.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the system focuses on analyzing content of speech only, then the processing simplicity is maintained, but the ability to read user emotion and attitude is insufficient
Solution Approach 1:
The patent segments the audio signal analysis into distinct components: silence period detection, filler sound detection, and meaningful speech detection. By dividing the audio stream into these separate analyzable segments, the system can extract multiple types of information (timing, emotion, attitude) from different portions of the speech without requiring overly complex unified processing
Solution Approach 2:
The patent adds a temporal dimension to speech analysis by examining not just the content but also the timing characteristics (silence periods between questions and answers, filler sounds during thinking). This multi-dimensional approach enables emotion and attitude recognition while maintaining processing feasibility through structured temporal analysis
2Loss of information
If the system processes all audio input uniformly, then processing simplicity is maintained, but the ability to capture thinking time and emotional cues is lost
Solution Approach 1:
The patent extracts specific information elements from the audio signal that are relevant to emotion and attitude recognition: silence period duration, filler sound presence, and meaningful speech content. By selectively extracting only these critical elements rather than processing the entire audio signal uniformly, the system preserves important information while avoiding unnecessary processing complexity
Solution Approach 2:
The patent introduces an intermediary processing layer that detects and identifies specific audio characteristics (silence periods, filler sounds) before final analysis. This intermediary detection step acts as a mediator between raw audio input and emotion/attitude interpretation, enabling sophisticated information extraction through structured intermediate processing
3Ease of operation
If the system only analyzes speech content, then processing speed is maintained, but the naturalness of interaction deteriorates
Solution Approach 1:
The patent performs preliminary detection of silence periods and filler sounds during the audio input phase itself, before full speech analysis begins. By pre-identifying these temporal characteristics, the system prepares emotion and attitude evaluation data in advance, enabling faster overall processing while maintaining natural interaction through comprehensive analysis
Data Source
AI summary
An information processing apparatus identifies, by using an audio signal acquired by collecting a user's voice, evaluation target time that includes at least either time not including the user's voice or time during which the user is producing a meaningless utterance and produces an output appropriate to the identified evaluation target time.


