Natural Language Camera Control via Visual Token Detection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computing devices that rely on graphical user interfaces and presence-sensitive technology for camera control can be cumbersome and prone to errors, especially when trying to capture moving objects, as they require manual inputs that may cause the device to move, blurring or affecting the quality of the photo or video.
Innovation Solution
A method and system that allow a computing device to receive natural language user inputs to control the camera, determining visual tokens to be captured and locating them within an image preview, enabling the device to capture images without requiring manual touch inputs, allowing for precise control and stabilization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a graphical user interface is used for camera control, then the device can receive user inputs to operate the camera, but the user may be too slow to provide inputs and the device may move causing blurred photos
Solution Approach 1:
The patent replaces the mechanical interaction system (touchscreen GUI requiring physical finger movements) with an acoustic field-based system (voice recognition). The user speaks natural language commands that are captured by the microphone and processed by the processor, eliminating the need for manual touchscreen interactions and associated device movement.
Solution Approach 2:
The patent introduces voice commands as an intermediary between the user and the camera control system. Instead of direct touchscreen interaction, the user's intent is conveyed through spoken language, which the system interprets and executes, providing a faster and more stable control method.
2Ease of operation
If manual touch inputs are used to control the camera, then the user can adjust settings and capture images, but the device movement during input causes blurred photos
Solution Approach 1:
The patent substitutes the mechanical touchscreen input system with an acoustic-based voice recognition system. This eliminates the physical contact and device movement associated with manual touch inputs, thereby preventing image blur while maintaining ease of operation.
Solution Approach 2:
The system allows the user to control the camera while keeping their hands steady on the device. The voice recognition system processes commands without requiring manual manipulation, enabling the user to maintain device stability for sharper photos.
3Ease of operation
If a graphical user interface is used for camera control, then the device can receive user inputs, but the interaction is cumbersome and impractical when framing the scene
Solution Approach 1:
The patent replaces the complex graphical user interface requiring multiple touchscreen interactions with a simple voice-based natural language system. This reduces the complexity of the control interface by eliminating the need for users to navigate through graphical menus and buttons while framing the scene.
Solution Approach 2:
Instead of requiring the user to manually navigate through a graphical interface to select capture parameters, the system inverts the approach by having the user speak their intent naturally and the system automatically interpreting and executing the appropriate camera controls based on the voice command.
Data Source
AI summary
In general, techniques of this disclosure may enable a computing device to capture one or more images based on a natural language user input. The computing device, while operating in an image capture mode, receive an indication of a natural language user input associated with an image capture command. The computing device determines, based on the image capture command, a visual token to be included in one or more images to be captured by the camera. The computing device locates the visual token within an image preview output by the computing device while operating in the image capture mode. The computing device captures one or more images of the visual token.


