Gesture-Controlled Directional Audio Recording for Group Conversations

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional methods for recording audio content in group conversations often result in poor audio quality and require significant human effort for transcription, as they either rely on manual transcription or omni-directional recording that captures all sounds simultaneously.

Innovation Solution

A system comprising a processor, image capturing device, audio recording device, and data storage that continuously captures images of participants, detects specific gestures to activate unidirectional audio recording towards the speaker, and stores audio data with timestamps, improving audio quality and reducing human transcription effort.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If an omni-directional microphone is used to record audio content in a group conversation, then all sounds are captured simultaneously, but the audio quality of individual speakers deteriorates

Engineering Contradiction:
Improvecoverage of audio recordingVSAvoidaudio quality
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The system dynamically switches between omni-directional and unidirectional recording modes based on real-time gesture detection. When a participant makes a specific gesture indicating they wish to speak, the system transitions to unidirectional recording focused on that participant, thereby maintaining high audio quality while preserving the ability to capture all participants when needed

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If manual transcription is used to record audio content, then human effort is required for transcription, but the accuracy of conversation record is improved

Engineering Contradiction:
Improveaccuracy of conversation recordVSAvoidtime for transcription
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system replaces the mechanical process of manual transcription with automated speech recognition technology. By capturing high-quality unidirectional audio and processing it through speech recognition algorithms, the system automatically generates conversation records, eliminating the need for manual transcription while maintaining accuracy and significantly reducing time consumption

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If unidirectional audio collection is performed toward a specific member, then audio quality of that member is improved, but the ability to capture other members deteriorates

Engineering Contradiction:
Improveaudio quality of specific speakerVSAvoidcapability to record multiple speakers
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The system employs periodic action by alternating between omni-directional recording (when no specific speaker is identified) and unidirectional recording (when a gesture is detected). This periodic switching allows the system to maintain versatility in capturing multiple speakers while periodically achieving high-quality focused recording when needed

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The recording system is made dynamic by automatically adjusting its directional characteristics based on real-time gesture detection. The system can transition from capturing all participants to focusing on a specific speaker and back again, thereby maintaining both adaptability and audio quality through dynamic reconfiguration

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11488596B2Method and system for recording audio content in a group conversation
Publication Date: 2022.11.01 CHEN HSIAO HAN
  • US11488596B2 patent drawing
  • US11488596B2 patent drawing
  • US11488596B2 patent drawing

AI summary

A method for recording audio content in a group conversation among a plurality of members includes: controlling an image capturing device to continuously capture images of the members; executing an image processing procedure on the images of the members to determine whether a specific gesture is detected; when the determination is affirmative, controlling an audio recording device to activate and perform directional audio collection with respect to a direction that is associated with the specific gesture to record audio data; and controlling a data storage to store the audio data and a time stamp associated with the audio data as an entry of conversation record.