Smart Wearable Inter-person Conversation Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for detecting inter-person conversations, such as face-to-face conversations, are inadequate as they struggle to differentiate between talking and listening, and require multiple devices or uncomfortable sensors, leading to inefficiencies and biases in data collection.
Innovation Solution
A smart wearable device captures audio data, extracts acoustic features, and uses neural network models to fuse neural network embedding features for effective detection of inter-person conversations by distinguishing between different speech patterns and speaker changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If multiple audio recording devices are used to capture conversation signals, then the ability to detect inter-person conversations improves, but the device complexity and ease of operation deteriorate
Solution Approach 1:
The patent combines multiple detection functions (speech detection, speaker change detection, and conversation detection) into a single integrated system using one wearable device. This merging of functions allows the system to achieve conversation detection capability without requiring multiple separate devices, thus resolving the contradiction between detection accuracy and device complexity.
Solution Approach 2:
The wearable device is designed to perform multiple functions: capturing audio signals, detecting speech presence, identifying speaker changes, and determining conversations. This multi-functionality allows a single device to replace what would traditionally require multiple specialized devices, improving ease of operation while maintaining detection precision.
2Measurement precision
If customized sensor boards are attached to detect respirational signals, then the ability to detect inter-person conversations improves, but the ease of operation and user comfort deteriorate
Solution Approach 1:
The system uses the wearable device's existing audio recording capabilities to detect conversations, eliminating the need for additional customized sensor boards. The device serves itself by using its built-in microphone and processing capabilities to perform conversation detection, thus maintaining user comfort while achieving detection accuracy.
Solution Approach 2:
The patent replaces the mechanical approach of attaching customized sensor boards to detect respirational signals with an acoustic approach using the device's built-in microphone. This substitution eliminates the need for uncomfortable physical attachments while maintaining the ability to detect conversations through audio signal analysis.
3Ease of operation
If self-report methods are used to document social interactions, then the ease of operation improves, but the measurement precision and objectivity deteriorate
Solution Approach 1:
The system automatically detects and documents conversations without requiring user input or self-reporting. The wearable device independently performs speech detection, speaker change detection, and conversation characterization, providing objective data while maintaining ease of operation since no user action is required beyond wearing the device.
Solution Approach 2:
The system provides automatic feedback about detected conversations, eliminating the need for manual self-reporting. Through continuous audio monitoring and analysis, the system objectively records social interaction data, resolving the contradiction between operational simplicity and measurement accuracy by making the detection process transparent and automatic.
Data Source
AI summary
A method, smart wearable device and computer program product for detecting inter-person conversations. Audio data is captured on the smart wearable device. Such audio data may be from various sources, including sounds that only involve listening by an individual (e.g., television show) and sounds that involves an individual participating in face-to-face communication. Acoustic features are then extracted from the captured audio data. Such extracted acoustic features are a description of the captured audio data. Neural network embedding features are then extracted from the extracted acoustic features using a first neural network model. The extracted neural network embedding features are then fused into a second neural network model (configured to detect inter-person conversations based on detecting speaker change point(s)) to perform user conversation inference. In this manner, inter-person conversations are more effectively detected using a smart wearable device.


