Voice Dialogue Target Detection via Image Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing devices struggle to accurately distinguish between a real live person and other objects that output voice, leading to unintended responses and reduced user satisfaction, particularly in environments with moving images and pets.

Innovation Solution

An information processing device with a determination unit that recognizes input images to identify dialogue targets based on voice output capability, and a dialogue function unit that controls voice interactions accordingly, preventing unnecessary responses.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If the device responds to any object that outputs voice, then the device can interact with more targets, but false detection occurs when the target is not a real live person (e.g., moving images on television)

Engineering Contradiction:
Improvevoice interaction capabilityVSAvoiddetection accuracy
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The patent introduces an image recognition unit as an intermediary between the voice input and the dialogue response system. This unit analyzes the visual characteristics of the voice source to determine if it is a real live person, thereby mediating between the broad voice interaction capability and the need for accurate detection to prevent false responses to moving images or other non-human sources.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the device performs detailed recognition to distinguish real persons from other objects, then detection accuracy improves, but system complexity increases

Engineering Contradiction:
Improvedialogue target detection precisionVSAvoidsystem structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges the image recognition unit with the existing voice processing system, combining visual and auditory analysis into a unified dialogue target detection system. This integration allows the system to achieve high detection precision by analyzing both image and voice characteristics simultaneously, while avoiding the complexity of completely separate recognition systems.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The image recognition unit serves multiple functions: it identifies whether the voice source is a real live person, determines the type of object emitting voice, and provides visual context for dialogue processing. This multi-functionality allows the system to achieve high detection precision without adding excessive complexity, as a single unit performs multiple recognition tasks.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11620997B2Information processing device and information processing method
Publication Date: 2023.04.04 SONY GROUP CORP
  • US11620997B2 patent drawing
  • US11620997B2 patent drawing
  • US11620997B2 patent drawing

AI summary

Provided is an information processing device that includes a determination unit that determines whether an object that outputs voice is a dialogue target related to voice dialogue based on a result of recognition of an input image, and a dialogue function unit that performs control related to the voice dialogue based on the determination. The dialogue function unit provides a voice dialogue function to the object based on the determination that the object being the dialogue target. Further provided is a method that includes determining whether an object that outputs voice is a dialogue target related to voice dialogue based on a result of recognition of an input image, and performing control related to the voice dialogue based on a result of the determining. The performing of the control further includes providing a voice dialogue function to the object based on the determination that the object is the dialogue target.