Image Processing Device for Accurate Remote Control Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing remote control systems struggle to accurately distinguish between multiple electronic devices of the same type, leading to potential incorrect device selection and operation, especially when devices are closely positioned or lined up along the user's line of sight.

Innovation Solution

An image processing device that generates an environment map using input images and feature data to identify and select candidate objects for remote control, including a data storage unit for object identification and feature data, an environment map storage unit, and a selecting unit that determines operable objects based on this data, allowing users to intuitively select the target device through a see-through type display.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech recognition technology is used to identify remote control target devices, then the system can recognize devices from user speech, but the recognition process becomes complicated when there are multiple same type devices

Engineering Contradiction:
Improvedevice recognition capabilityVSAvoidrecognition process complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the device identification process into multiple components: speech recognition for device type, image recognition for device location and appearance, and environment map for spatial positioning. This segmentation allows each component to handle a specific aspect, reducing overall complexity when multiple devices are present

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an environment map as an intermediary data structure that stores pre-recognized device information, positions, and relationships. This intermediary allows the system to resolve ambiguities between multiple devices by cross-referencing speech input with stored environmental data, reducing recognition process complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If only gesture recognition is used to identify target devices, then the system can operate devices through gestures, but it is difficult to distinguish between multiple devices located at positions along the user's line of sight

Engineering Contradiction:
Improvegesture-based controlVSAvoiddevice selection accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent merges multiple recognition methods: gesture recognition for operational intent, image recognition for device identification, and environment map for position verification. This combination maintains ease of gesture-based operation while improving device selection accuracy through multi-modal verification

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent adds spatial dimension to device selection by using environment maps that store three-dimensional positions and relationships of devices. This allows the system to distinguish between devices along the line of sight by referencing their precise spatial coordinates, transforming a two-dimensional gesture problem into a three-dimensional solution

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Quantity of substance

If multiple electronic devices are positioned closely or lined up, then the environment can accommodate more devices, but accurate identification and selection of the intended device becomes difficult

Engineering Contradiction:
Improvenumber of devicesVSAvoiddevice identification accuracy
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent performs preliminary actions by pre-building environment maps that contain detailed information about device positions, appearances, and relationships before user interaction. This advance preparation allows the system to quickly and accurately identify the intended device even when multiple devices are closely positioned, maintaining identification accuracy regardless of device quantity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent uses visual differentiation through image recognition data, which can detect and distinguish devices based on their visual characteristics such as color, shape, and size. This allows users to identify and select the correct device among closely positioned devices by visual cues displayed on the terminal device screen

Inventive Principle:
Principle #32Color changes

4Measurement precision

If speech modifiers are required to specify target devices, then accurate device selection is possible, but the operation becomes more complex and requires additional user input

Engineering Contradiction:
Improvedevice specification accuracyVSAvoidoperation simplicity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent implements feedback by displaying recognized device information on the terminal device screen before execution. This allows users to verify device identification accuracy without requiring modifiers, and to correct selections if needed, maintaining both accuracy and operational simplicity through visual confirmation

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS10908676B2Image processing device, object selection method and program
Publication Date: 2021.02.02 SONY GROUP CORP
  • US10908676B2 patent drawing
  • US10908676B2 patent drawing
  • US10908676B2 patent drawing

AI summary

There is provided an image processing device including: a data storage unit that stores object identification data for identifying an object operable by a user and feature data indicating a feature of appearance of each object; an environment map storage unit that stores an environment map representing a position of one or more objects existing in a real space and generated based on an input image obtained by imaging the real space using an imaging device and the feature data stored in the data storage unit; and a selecting unit that selects at least one object recognized as being operable based on the object identification data, out of the objects included in the environment map stored in the environment map storage unit, as a candidate object being a possible operation target by a user.