Speech Recognition Apparatus with Adaptive Acoustic Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems for digital cameras face accuracy deterioration when state information of movable or connected devices changes, such as lens settings or microphone configurations, leading to suboptimal recognition of user commands.

Innovation Solution

A speech recognition apparatus and method that acquires state information of movable and connected devices, adjusts the control content for speech recognition based on this information, and outputs command signals to operate the digital camera, using multiple microphones and adaptive acoustic models to improve recognition accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If state information of movable portions or connected devices is changed, then the adaptability of the speech recognition system is improved, but the recognition accuracy is deteriorated

Engineering Contradiction:
Improveadaptability of speech recognition systemVSAvoidrecognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent implements dynamic adjustment of speech recognition parameters based on real-time state information of movable portions (lens, display, illumination) and connected devices (microphone, speaker). The recognition control portion continuously adapts acoustic models and processing parameters according to the current device state, transforming the static recognition system into a dynamic one that responds to environmental changes while maintaining accuracy through state-aware processing.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes physical and processing parameters of the speech recognition system based on device state information. Specifically, it adjusts acoustic models, signal processing parameters, and recognition thresholds according to the positions of movable portions and characteristics of connected devices. This parameter adaptation resolves the contradiction by making the system both adaptable to different states and accurate within each state through optimized parameters.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If multiple microphones and adaptive acoustic models are used, then the speech recognition accuracy is improved, but the device complexity is increased

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidcomplexity of speech recognition apparatus
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The recognition control portion serves multiple functions: it manages state information acquisition, adjusts acoustic models, processes speech signals from multiple microphones, and controls output to connected devices. By making this single control portion multi-functional, the patent achieves high recognition accuracy through coordinated processing without proportionally increasing overall system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system implements feedback loops where state information from movable portions and connected devices is continuously acquired and used to adjust recognition parameters. The recognition control portion monitors device state changes and dynamically modifies acoustic models and processing parameters, creating a closed-loop system that maintains accuracy while managing complexity through intelligent feedback-based adaptation rather than brute-force hardware increases.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20240331693A1Speech recognition apparatus, speech recognition method, speech recognition program, and imaging apparatus
Publication Date: 2024.10.03 NIKON CORP
  • US20240331693A1 patent drawing
  • US20240331693A1 patent drawing
  • US20240331693A1 patent drawing

AI summary

A speech recognition apparatus includes an acquisition portion that is configured to acquire state information regarding at least one of a movable portion in a target device operated according to an input speech or a connected device connected to the target device; a recognition control portion that is configured to set a control content for recognizing the speech based on the state information acquired by the acquisition portion and to recognize the speech; and an output portion that is configured to output a command signal for operating the target device to the target device according to a recognition result of the recognition control portion.