Voice-Controlled Box Management in Image Processing Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing image processing systems with box functions, used in multi-function peripherals, face usability issues due to complex operations required for users to access and manage image data stored in multiple boxes, especially when users forget where they saved documents.
Innovation Solution
An image processing system that utilizes a microphone to obtain voice input from users, processes this input to specify the correct box among multiple boxes based on user identifiers, and informs the user about the contents of the specified box, thereby simplifying the access process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If users access image data by checking box contents on the UI, then the user can find the desired document, but the operation becomes complicated and usability degrades
Solution Approach 1:
The patent replaces the mechanical interaction of manually checking box contents on the UI with voice-based interaction. The interaction agent processes natural language voice commands to identify and retrieve documents, substituting the manual browsing mechanism with an automated voice recognition system that directly accesses and presents document information without requiring users to navigate through box contents visually.
2Quantity of substance
If multiple boxes are provided for storing image data, then the storage capacity and organization improve, but the user cannot remember which box contains which document
Solution Approach 1:
The interaction agent provides feedback by listening to user voice descriptions of documents and responding with the location or contents of the specified box. When a user asks about a document, the system analyzes the voice input, searches the appropriate boxes, and provides feedback information about where the document is stored or what contents are in a particular box, creating a closed-loop interaction that helps users track document locations without manually remembering them.
Solution Approach 2:
The interaction agent serves as an intermediary between the user and the multiple storage boxes. Instead of users directly managing and remembering the contents of multiple boxes, the interaction agent mediates by processing voice commands, searching through boxes based on document descriptions, and presenting relevant information, thereby shielding users from the complexity of tracking multiple storage locations.
3Ease of operation
If voice recognition is implemented for box operations, then user convenience improves, but the system cannot accurately specify the correct box when multiple boxes exist
Solution Approach 1:
The system segments the voice recognition process into distinct analytical components: extracting document description features from voice input, comparing these features against stored document metadata, and using the comparison results to identify the specific box containing the target document. This segmentation allows the system to systematically process voice commands and accurately distinguish between multiple boxes by matching specific document characteristics rather than treating all boxes uniformly.
Data Source
AI summary
An image processing system capable of managing image data using a plurality of boxes, comprises a microphone that obtains a sound, an obtaining unit that obtains a user identifier based on voice information of a user obtained via the microphone, a specifying unit that specifies one box among the plurality of boxes based on specification information including at least the user identifier, and an informing unit that informs the user of information related to the specified one box.


