AR Image Processing Using Sound Recognition for Virtual Object Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional Augmented Reality (AR) systems lack interactiveness and operability due to the inability to perform operations on virtual objects displayed in real-world images without physically interacting with the imaging device, limiting user input to only moving the camera, which restricts user engagement and interaction.

Innovation Solution

An image processing program that uses sound recognition to control the position, orientation, and display form of virtual objects in a real-world image, allowing users to interact with virtual objects using voice commands, enhancing input methods and user engagement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If conventional camera-based input methods are used in AR systems, then the system structure remains simple, but user interactiveness and operability deteriorate due to limited input methods

Engineering Contradiction:
Improveuser interactivenessVSAvoidsystem structure
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The patent replaces mechanical camera-based input methods with acoustic sound recognition methods. Instead of requiring users to physically manipulate cameras or wear imaging devices, the system uses sound input devices (microphones) to capture voice commands, which are then processed by sound recognition means to control virtual objects. This substitution significantly improves ease of operation while adding minimal complexity to the system structure.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Ease of operation

If users physically interact with imaging apparatus to perform input operations, then input precision may improve, but user operability deteriorates due to physical restrictions while taking images

Engineering Contradiction:
Improveuser operabilityVSAvoidinput stability
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent substitutes mechanical physical interaction with acoustic field-based interaction. Users can issue voice commands while maintaining stable camera positioning, eliminating the conflict between physical manipulation and image stability. The sound recognition means accurately captures and processes voice inputs without being affected by camera movement or physical handling of the imaging device.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Adaptability or versatility

If only camera movement is used for user input, then system complexity remains low, but interactiveness and user engagement deteriorate due to limited input variation

Engineering Contradiction:
Improveinput method variationVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements multi-functionality by integrating sound recognition capabilities into the existing AR system. The sound input device and sound recognition means can recognize various types of sounds (voice commands, clapping, whistling) and translate them into different input operations, enabling diverse interactiveness without requiring separate devices for each input method. This universal approach significantly increases input method variation while adding moderate system complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS8854356B2Storage medium having stored therein image processing program, image processing apparatus, image processing system, and image processing method
Publication Date: 2014.10.07 NINTENDO CO LTD
  • US8854356B2 patent drawing
  • US8854356B2 patent drawing
  • US8854356B2 patent drawing

AI summary

An image taken by a real camera is repeatedly obtained, and position and orientation information determined in accordance with a position and an orientation of a real camera in a real space is repeatedly calculated. A virtual object or a letter to be additionally displayed on the taken image is set as an additional display object, and based on a result of recognition of a sound inputted into a sound input device, at least one selected from the group consisting of a display position, an orientation, and a display form of the additional display object is set. A combined image repeatedly generated by superimposing on the taken image the set additional display object with reference to a position in the taken image in accordance with the position and orientation information is displayed on a display device.