Earphone Conversation Detection for Accurate Volume Adjustment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing conversation awareness features in earphones inaccurately activate or fail to detect conversations, particularly in noisy environments and during user activities like singing or humming, leading to inconsistent volume adjustments and disrupted user experiences.

Innovation Solution

A method and system utilizing a neural network model to analyze audio signals, user activity, and environmental factors to accurately detect conversations by determining relatedness, intelligibility, and directional probabilities, adjusting playback volume accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conversation awareness feature activates on any speech detection, then user can converse while listening to media, but media volume is mistakenly reduced during non-conversation activities like singing or humming

Engineering Contradiction:
Improveconversation awareness activationVSAvoidaccurate conversation detection
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system dynamically adjusts conversation detection thresholds and parameters based on the detected activity state. When singing or humming is detected through audio pattern analysis, the system transitions to a different detection mode that prevents false conversation awareness activation, while maintaining normal detection during actual conversations.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes detection parameters such as voice activity thresholds, frequency range sensitivity, and temporal patterns based on the current activity state. By modifying these parameters dynamically, the system distinguishes between conversational speech and non-conversational vocal activities like singing or humming.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If conversation awareness lowers media volume on speech detection, then ambient voices are amplified, but volume is reduced even when no conversation is occurring

Engineering Contradiction:
Improveautomatic volume adjustmentVSAvoidspeech detection accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system implements feedback mechanisms where the detected activity state continuously informs the volume control decisions. When non-conversation activities are detected, the feedback loop prevents volume reduction, while actual conversations trigger appropriate volume adjustments through the conversation awareness feature.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary activity state classification before triggering volume adjustments. By analyzing audio patterns in advance to determine whether the detected speech constitutes a actual conversation or non-conversational activity, the system prevents premature or incorrect volume changes.

Inventive Principle:
Principle #10Preliminary action

3Device complexity

If earphones detect user voice only, then conversation detection is simple, but feature fails to detect when others address the user in noisy environments

Engineering Contradiction:
Improvedetection systemVSAvoidconversation detection reliability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The audio detection system is segmented into multiple independent detection channels: one for user voice detection and another for ambient speech detection. This segmentation allows the system to monitor both the user's own voice and external speech sources simultaneously, enabling reliable conversation detection even in noisy environments where others are addressing the user.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The detection system is designed with multi-functionality to handle multiple detection scenarios: user self-speech, external speech addressing the user, and ambient noise. By making the detection system universal across these different functions, it reliably detects conversations regardless of who is speaking or the noise level.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Adaptability or versatility

If conversation awareness activates on any vocal activity, then user can be aware of surroundings, but unintended activation occurs during singing, humming, or lip-syncing

Engineering Contradiction:
Improveconversation awareness featureVSAvoidunintended volume disruption
Core Design Contradiction:
Adaptability or versatilityVSObject-generated harmful factors

Solution Approach 1:

The system employs periodic analysis of audio patterns to distinguish between conversational speech and non-conversational vocal activities. By analyzing the temporal periodicity, rhythm, and pattern characteristics of detected vocal activities, the system can identify singing or humming patterns and prevent false conversation awareness activation during these periodic non-conversational activities.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS20260082156A1Method and system for intelligent conversation detection in earphones
Publication Date: 2026.03.19 SAMSUNG ELECTRONICS CO LTD
  • US20260082156A1 patent drawing
  • US20260082156A1 patent drawing
  • US20260082156A1 patent drawing

AI summary

A method for conversation detection in earphones includes detecting an audio signal by at least one microphone of the earphones, wherein the detected audio signal includes a user voice; computing a relatedness score of the user voice with a playback audio of a user device; determining an intelligibility score of the user voice for a predetermined distance; determining a situation context of the user of the earphones in the detected audio signal; determining a directional probability of a conversation based on at least sensor data and the situation context; determining, using a neural network model, a conversation probability based on at least the relatedness score, the intelligibility score, the situation context, and the directional probability; adjusting a volume of the playback audio based on at least the conversation probability; and restoring the volume of the playback audio in the earphones based on determining an end of the conversation.