Multi-Stage Voice Command Search for MFP Screens

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image processing devices, such as multi-functional peripherals (MFPs), cannot properly detect voice operation commands intended for operation screens other than the currently displayed screen due to a fixed search range for voice recognition data, limiting the ability to execute instructions for buttons on secondary or caller screens.

Innovation Solution

An image processing device with a display means, an obtain means for voice recognition data, a determine means to identify a search target character string, a search means that executes two-stage search processing across priority-ordered voice operation command groups related to both the current and caller screens, allowing detection of voice operation commands for multiple screens.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the search range is fixed to only the current layer screen's voice operation command group, then the processing is simple and fast, but voice operation commands related to other layer screens cannot be detected

Engineering Contradiction:
Improvevoice operation command detection capabilityVSAvoidsearch processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the search processing into two distinct stages: first search processing that searches only the current layer screen's voice operation command group, and second search processing that searches the caller layer screen's voice operation command group. This segmentation allows the system to maintain simplicity in the primary search while adding extended functionality through the optional second search stage, thereby resolving the contradiction between detection capability and processing complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements dynamic search range adjustment based on the detected voice operation command. The system first performs a limited search in the current layer, and only if no matching command is found, does it dynamically expand the search range to include the caller layer screen's command group. This dynamic approach allows the system to adapt its complexity to the actual needs of each voice input scenario.

Inventive Principle:
Principle #15Dynamics

2Reliability

If the search range is expanded to include both current and caller screens, then all voice operation commands can be detected, but the search processing time and complexity increase

Engineering Contradiction:
Improvevoice operation command detection accuracyVSAvoidvoice recognition processing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

By dividing the search into two sequential stages with different scopes, the patent ensures that most common commands (those on the current layer) are processed quickly, while less common commands (on the caller layer) receive extended search coverage. This segmentation maintains high detection accuracy for all commands while minimizing average processing time.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies partial action by performing a complete search only when necessary (when the first search yields no results). In the majority of cases where the voice command corresponds to the current layer screen, only the partial first search is executed, thereby reducing average processing time while maintaining reliability for all possible commands.

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If multi-stage search processing is implemented, then voice operation commands on caller screens can be detected, but the device complexity increases

Engineering Contradiction:
Improvecross-screen voice command supportVSAvoidsearch means complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the search functionality into modular components: a first search means for current layer commands and a second search means for caller layer commands. This modular segmentation allows the system to implement complex multi-screen search capability while maintaining clear separation of functions, making the overall system more manageable and less complex than a monolithic approach would be.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP3528244B1Image processing device, method for controlling image processing device, and program
Publication Date: 2021.09.29 KONICA MINOLTA INC
  • EP3528244B1 patent drawingFigure 1
  • EP3528244B1 patent drawingFigure 2
  • EP3528244B1 patent drawingFigure 3

AI summary

Provide is a technology that enables to properly detect one voice operation command corresponding to user's voice input from among a plurality of voice operation commands related to a plurality of operation screens (210, 230). The MFP (10) obtains a voice recognition result (voice recognition data) related to a voice vocalized in a state in which the operation screens (210, 230) are displayed on the touch panel (45), and determines a search target character string. The MFP (10) first executes the first search processing in which a search range is the first command group (M1) (630 and the like) to which the first priority order has been given among a plurality of voice operation commands including: the voice operation command group (610) related to the screen (210); and the voice operation command group (630) related to the screen (230) displayed according to user's operation for the screen (210). In a case where the search target character string has not been detected by the first search processing, the second search processing in which a search range is the second command group (M2) (610 and the like) to which second priority order has been given among the plurality of voice operation commands is executed.