Autonomous Unit State Transition for Speech Recognition Noise

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face challenges such as reduced accuracy due to operating noise from autonomous devices, complex user interactions, time-consuming wake-up word recognition, difficulty in natural conversation, and limitations in expressing speech transmission methods.

Innovation Solution

An information processing apparatus and method that controls an autonomous operation unit to transition through multiple states based on detected triggers, including a first active state where autonomous actions are restricted and a second active state for speech recognition, allowing continuous sound streaming and processing, and using environmental and user-specific triggers to manage these states.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If the autonomous operation unit performs autonomous actions continuously, then productivity is improved, but speech recognition accuracy deteriorates due to operating noise

Engineering Contradiction:
Improveautonomous action executionVSAvoidspeech recognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system dynamically transitions between different operational states (first active state with restricted actions and second active state with speech recognition). The control unit adjusts the autonomous operation unit's behavior based on detected triggers, allowing the system to adapt its characteristics in real-time to resolve the contradiction between continuous operation and speech recognition accuracy

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements periodic alternation between autonomous action execution and speech recognition modes. By controlling the autonomous operation unit to restrict actions during speech recognition periods and resume actions during other periods, the system achieves both continuous productivity and accurate speech recognition through time-division multiplexing

Inventive Principle:
Principle #19Periodic action

2Measurement precision

If wake-up word recognition is required for each interaction, then speech recognition accuracy is improved, but loss of time increases due to repeated wake-up sequences

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidinteraction response time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary action by detecting triggers and transitioning to the second active state (speech recognition mode) before the user actually speaks. This anticipatory state transition eliminates the need for repeated wake-up words, as the system is already in readiness to recognize speech when triggered by environmental cues or user presence detection

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system maintains continuous speech recognition capability by staying in the second active state for extended periods after trigger detection. This continuous readiness state eliminates interruptions caused by wake-up word requirements, allowing seamless conversation flow while maintaining accurate speech recognition through sustained microphon e activation

Inventive Principle:
Principle #20Continuity of useful action

3Measurement precision

If the autonomous operation unit restricts actions to reduce noise, then speech recognition accuracy is improved, but productivity decreases due to limited autonomous operations

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidautonomous operation capability
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system dynamically adjusts its operational characteristics by transitioning between states based on trigger detection. During the first active state, autonomous actions are restricted to minimize noise. When triggered, the system transitions to the second active state where speech recognition is prioritized. After speech processing, the system can return to autonomous operations, creating a dynamic balance between noise reduction and productivity

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements periodic cycles of autonomous operation followed by speech recognition modes. During speech recognition periods, actions are restricted to ensure accuracy. Between these periods, the system resumes autonomous actions to maintain productivity. This periodic alternation ensures both speech recognition accuracy and continuous operational capability are achieved over time

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS12014736B2Information processing apparatus and information processing method
Publication Date: 2024.06.18 SONY GROUP CORP
  • US12014736B2 patent drawing
  • US12014736B2 patent drawing
  • US12014736B2 patent drawing

AI summary

An information processing apparatus that includes a control unit controlling an action of an autonomous operation unit, and in which the control unit controls transition of plural states relating to speech recognition processing through the autonomous operation unit based on a detected trigger, and the states include a first active state in which an action of the autonomous operation unit is restricted, and a second active state in which the speech recognition processing is performed. An information processing method in which a processor controls an action of an autonomous operation unit, the controlling includes controlling transition of plural states relating to speech recognition processing through the autonomous operation unit based on a detected trigger, and the states include a first active state in which an action of the autonomous operation unit is restricted, and a second active state in which the speech recognition processing is performed.