Semantic Interpretation Unit for Voice Command Context

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice command systems for devices with autonomous movement, such as cars and robots, struggle to accurately interpret voice recognition results due to lack of consideration for the context in which the voice is collected, leading to misinterpretation of commands, especially those involving relative directions like 'right' or 'left'.

Innovation Solution

An information processing apparatus and method that incorporates a semantic interpretation unit to interpret voice recognition results based on both the recognition result and context information, including positional relationships and image display orientations, to accurately determine the user's intended command.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voice recognition is performed without considering context information, then the processing speed is fast and the system is simple, but the interpretation accuracy deteriorates and misinterpretation occurs

Engineering Contradiction:
Improvevoice interpretation accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary action by collecting context information (image data, device state, positional relationships) before interpreting the voice command. This allows the semantic interpretation unit to have relevant contextual data ready when voice recognition occurs, improving interpretation accuracy without adding complexity during the critical voice processing moment.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The semantic interpretation unit acts as an intermediary between voice recognition and command execution. It receives both the recognition result and context information, integrates them, and produces the final interpretation. This mediator structure allows complex contextual processing while keeping the voice recognition and execution paths relatively simple.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If context information is incorporated into voice interpretation, then the interpretation accuracy improves, but the processing time increases and productivity decreases

Engineering Contradiction:
Improvevoice interpretation accuracyVSAvoidcommand processing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

Context information such as image data, device state, and positional relationships is collected and prepared in advance before voice interpretation occurs. This preliminary preparation ensures that when the voice command needs interpretation, the contextual data is already available, reducing the actual processing time during voice interaction.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The semantic interpretation unit autonomously determines which context information is relevant and integrates it with the voice recognition result without requiring external intervention or complex coordination. This self-service approach streamlines the integration process and maintains processing efficiency.

Inventive Principle:
Principle #25Self-service

3Reliability

If semantic interpretation based on context is implemented, then the reliability of voice commands improves, but the device complexity increases

Engineering Contradiction:
Improvevoice command reliabilityVSAvoidprocessing system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The semantic interpretation unit is designed as a universal component that handles multiple types of context information (image data, device state, positional relationships, user preferences) through a single integrated process. This multi-functional design improves reliability across different scenarios without proportionally increasing system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The semantic interpretation unit serves as an intermediary layer that standardizes the integration of diverse context information types. By creating a unified interface between context data and voice interpretation, it manages complexity while ensuring reliable command execution across various situations.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10522145B2Information processing apparatus and information processing method
Publication Date: 2019.12.31 SONY GROUP CORP
  • US10522145B2 patent drawing
  • US10522145B2 patent drawing
  • US10522145B2 patent drawing

AI summary

An information processing apparatus and information processing method are provided to interpret the meaning of a result of voice recognition adaptively to the situation in collecting voice. The information processing apparatus includes a semantic interpretation unit that interprets a meaning of a recognition result of a collected voice of a user on a basis of the recognition result and context information in collecting the voice.