Speech Recognition Controller for Image Forming Systems

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image forming systems that utilize speech input are not user-friendly, as they struggle to process natural language and often require repeated attempts due to subtle accent differences or irrelevant words, limiting their ability to respond to user operations and screen states effectively.

Innovation Solution

An image forming system that includes a controller to process speech input in natural language, using a combination of speech recognition, morphological analysis, and machine learning to associate spoken words with image formation settings, allowing users to operate the system by specifying settings through spoken commands without the need for pre-defined keywords.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition with pre-registered keywords is used, then the system can identify specific commands, but it fails to process natural language variations and accent differences

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidnatural language processing capability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent introduces an intermediary processing layer between speech input and command execution. This layer includes morphological analysis and machine learning components that bridge the gap between raw speech signals and predefined commands, enabling the system to handle natural language variations while maintaining accurate command identification

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system performs preliminary processing of speech input through morphological analysis before matching with commands. This preliminary action breaks down speech into linguistic components, allowing the system to handle variations in wording, accent, and sentence structure while still identifying the intended command

Inventive Principle:
Principle #10Preliminary action

2Reliability

If fixed speech input mechanism is used, then the system can reliably execute registered commands, but it cannot respond to user operations and screen states

Engineering Contradiction:
Improvecommand execution reliabilityVSAvoiduser-friendly operation capability
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The patent transforms the static, fixed speech input mechanism into a dynamic system that adapts to current screen states and user operations. The speech recognition system now changes its behavior based on what is displayed and what the user is doing, making the interaction more intuitive and easier to use while maintaining reliable command execution

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates feedback mechanisms where the current screen state and user operations inform the speech recognition process. This feedback loop allows the system to understand context, disambiguate commands, and provide more accurate responses based on the current operational state, enhancing both reliability and ease of use

Inventive Principle:
Principle #23Feedback

3Measurement precision

If accent-sensitive speech recognition is used, then specific commands can be identified, but subtle accent differences cause mismatch and require repeated attempts

Engineering Contradiction:
Improvecommand identification precisionVSAvoidtime for repeated speech attempts
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies partial matching through morphological analysis, where the system analyzes linguistic components of speech rather than requiring exact word-for-word matches. This approach is sufficiently precise to identify commands while being tolerant of accent variations and wording differences, eliminating the need for repeated attempts

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11792338B2Image processing system for controlling an image forming apparatus with a microphone
Publication Date: 2023.10.17 CANON KK
  • US11792338B2 patent drawing
  • US11792338B2 patent drawing
  • US11792338B2 patent drawing

AI summary

An image forming system is configured to receive an input of natural language speech. Regardless of whether the natural language speech includes a combination of first words or second words, the image forming system can recognize the natural language speech as an instruction to select a specific print setting displayed on a screen.