Speech Control System Selective Response for Image Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Multifunctional image processing apparatuses using speech UI face challenges in reducing user burden due to recognition errors and redundant responses, which complicate user interaction and increase operational complexity.

Innovation Solution

A speech control system comprising a microphone, speaker, and controller that selectively outputs response speech based on the number of setting items, either reading all specified items or omitting at least one, to prevent redundant information and reduce user burden, using a reading condition to determine the appropriate response.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the system outputs all recognized setting items by speech, then the user can confirm the speech recognition results, but the response speech becomes redundant and increases user burden

Engineering Contradiction:
Improvespeech recognition accuracy confirmationVSAvoiduser burden
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent extracts only the necessary setting items for speech output based on a determination result, rather than outputting all recognized items. The controller selectively outputs setting items that need user confirmation, removing redundant information from the response speech.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by outputting only a subset of setting items (those determined to need confirmation) rather than all recognized items. This partial output approach reduces redundancy while maintaining necessary confirmation functionality.

Inventive Principle:
Principle #16Partial or excessive action

2Loss of information

If the system provides comprehensive speech feedback, then the user can verify all settings, but the interaction complexity increases

Engineering Contradiction:
Improvesetting information completenessVSAvoidinteraction complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent extracts and outputs only the specific setting items that require user confirmation, rather than providing comprehensive feedback on all settings. This selective extraction reduces interaction complexity while preserving essential information.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the setting items into two groups: those that need speech output for confirmation and those that don't. This segmentation allows the system to manage information delivery more efficiently, reducing overall interaction complexity.

Inventive Principle:
Principle #1Segmentation

3Reliability

If the system reads all setting items, then recognition errors can be detected, but the response time increases

Engineering Contradiction:
Improverecognition error detectionVSAvoidresponse time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent extracts only the setting items that need confirmation for speech output, reducing the total time required to deliver the response while maintaining error detection capability for critical items.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by confirming only necessary setting items through speech rather than all items, thereby reducing response time while maintaining sufficient reliability for error detection.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11475892B2Speech control system, speech control method, image processing apparatus, speech control apparatus, and storage medium
Publication Date: 2022.10.18 CANON KK
  • US11475892B2 patent drawing
  • US11475892B2 patent drawing
  • US11475892B2 patent drawing

AI summary

There is provided a speech control system including: a microphone configured to acquire speech; a speaker configured to output speech; an image processing unit; and a controller configured to control settings of the image processing unit. The controller is configured to: specify one or more setting items represented by an input speech of a user acquired by the microphone that are to be set for the image processing unit, and depending on whether or not the specified one or more setting items satisfy a reading condition, cause the speaker to output a first response speech that reads the one or more setting items, or a second response speech that does not read at least one out of the one or more setting items.