Voice-Controlled Camera Recording with Wake Word Timestamping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-controlled image pickup systems struggle to accurately distinguish and remove unnecessary control voices from recorded video, particularly when multiple control voices are present, leading to ambiguity in determining which voice initiates recording operations.
Innovation Solution
An image pickup apparatus equipped with an audio input unit and a control unit that utilizes wake words and control words to manage video recording, where the control unit stops recording and records video data prior to the start time of the wake word, effectively isolating and removing unnecessary control voices from the video file.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If voice recognition technology is used to control video recording operations, then hands-free operation convenience is improved, but ambiguity arises when multiple control voices are present making it unclear which voice should trigger recording
Solution Approach 1:
The system performs preliminary voice detection by identifying wake words before executing recording operations. When a wake word is detected, the system activates recording mode and stores the wake word's start time, ensuring that subsequent control voices are captured accurately without ambiguity about which voice should trigger operations.
Solution Approach 2:
The patent introduces wake words as intermediary signals between the user and the recording system. These wake words serve as explicit triggers that activate recording mode, creating a clear intermediate state that resolves the ambiguity of directly responding to any control voice without proper context.
2Speed
If recording starts immediately upon detecting any control voice, then response speed is improved, but unnecessary control voices are captured in the video recording
Solution Approach 1:
The system performs preliminary detection of wake words before activating recording. By storing the wake word start time and using it as a reference point, the system ensures recording only begins when appropriate, preventing capture of unnecessary control voices while maintaining fast response to valid triggers.
Solution Approach 2:
The patent extracts and isolates wake words from the audio stream as distinct trigger events. By separating wake word detection from general voice recognition, the system can identify specific recording trigger points and exclude unrelated control voices from the recording, preventing unnecessary data capture.
3Reliability
If the system records video data from the current time, then recording completeness is improved, but control voices are included in the recorded video
Solution Approach 1:
The system performs preliminary detection and storage of wake word start times before recording begins. By using this pre-detected timestamp as the recording start point, the system ensures completeness of relevant content while excluding control voices that occurred before the explicit recording trigger.
Solution Approach 2:
The patent extracts the wake word timestamp from the audio stream and uses it as a precise boundary marker. This extraction allows the system to separate useful video content from unwanted control voices by starting recording exactly at the wake word occurrence, maintaining completeness while removing contamination.
Data Source
AI summary
An image pickup apparatus includes an image pickup unit that obtains video, an audio input unit that collects sound and a control unit that controls recording of the video based on a wake word and a control word included in the sound collected by the audio input unit. In a case where the control word that gives an instruction to stop recording the video is included in the sound, the control unit stops recording the video and records video data before a start time of the wake word as a video file.


