Information Processing Device with Whisper-Based Command Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional speech recognition systems struggle to distinguish between text input and symbol or command input, requiring users to manually differentiate between normal voice and whisper, which complicates the input process and burdens the user with unnecessary cognitive load.

Innovation Solution

An information processing device that classifies voice input into normal voice and whisper using a first neural network, and processes normal voice for text input while using a second neural network to recognize whispers for symbol or command input, allowing seamless integration of both modes without additional hardware.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional speech recognition is used to input text and commands, then text input is achieved, but the system cannot distinguish between text input and symbol/command input, causing all inputs to be treated as text

Engineering Contradiction:
Improveinput mode differentiationVSAvoidinput type recognition accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent changes the acoustic parameter of voice input by introducing whisper mode as a distinct parameter from normal voice. The voice recognition system detects the whisper parameter to differentiate between text input (normal voice) and symbol/command input (whisper), enabling accurate input type recognition without additional hardware

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If users are required to manually differentiate between normal voice and whisper for different input types, then input accuracy is improved, but user burden and cognitive load increase

Engineering Contradiction:
Improveinput type recognition accuracyVSAvoiduser operation simplicity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system performs self-service by automatically detecting and classifying the input type based on the acoustic characteristics of the voice. The voice recognition unit autonomously determines whether the input is text or symbol/command based on whisper detection, eliminating the need for users to manually indicate input type and reducing cognitive load

Inventive Principle:
Principle #25Self-service

3Measurement precision

If additional hardware is introduced to enable whisper detection and classification, then whisper recognition accuracy is improved, but device complexity increases

Engineering Contradiction:
Improvewhisper detection accuracyVSAvoidsystem structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The existing voice recognition unit is made multi-functional by enabling it to perform both normal voice recognition and whisper detection. The same hardware component processes both input types by analyzing acoustic parameters, eliminating the need for separate dedicated hardware and maintaining system simplicity while achieving accurate whisper recognition

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250273203A1Information processing device, information processing method, and computer program
Publication Date: 2025.08.28 SONY GROUP CORP
  • US20250273203A1 patent drawing
  • US20250273203A1 patent drawing
  • US20250273203A1 patent drawing

AI summary

Provided is an information processing device that performs processing related to voice input.The information processing device includes: a classification unit that classifies an uttered voice into a normal voice and a whisper on the basis of a voice feature amount; a recognition unit that recognizes a whisper classified by the classification unit; and a control unit that controls processing based on a recognition result of the recognition unit. The information processing device further includes a normal voice recognition unit that recognizes a normal voice classified by the classification unit, in which the control unit performs processing corresponding to a recognition result of a whisper by the recognition unit on a recognition result of the normal voice recognition unit.