AR Image Processing Using Sound Recognition for Virtual Object Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional Augmented Reality (AR) systems lack interactiveness and operability due to the inability to perform operations on virtual objects displayed in real-world images without physically interacting with the imaging device, limiting user input to only moving the camera, which restricts user engagement and interaction.
Innovation Solution
An image processing program that uses sound recognition to control the position, orientation, and display form of virtual objects in a real-world image, allowing users to interact with virtual objects using voice commands, enhancing input methods and user engagement.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If conventional camera-based input methods are used in AR systems, then the system structure remains simple, but user interactiveness and operability deteriorate due to limited input methods
Solution Approach 1:
The patent replaces mechanical camera-based input methods with acoustic sound recognition methods. Instead of requiring users to physically manipulate cameras or wear imaging devices, the system uses sound input devices (microphones) to capture voice commands, which are then processed by sound recognition means to control virtual objects. This substitution significantly improves ease of operation while adding minimal complexity to the system structure.
2Ease of operation
If users physically interact with imaging apparatus to perform input operations, then input precision may improve, but user operability deteriorates due to physical restrictions while taking images
Solution Approach 1:
The patent substitutes mechanical physical interaction with acoustic field-based interaction. Users can issue voice commands while maintaining stable camera positioning, eliminating the conflict between physical manipulation and image stability. The sound recognition means accurately captures and processes voice inputs without being affected by camera movement or physical handling of the imaging device.
3Adaptability or versatility
If only camera movement is used for user input, then system complexity remains low, but interactiveness and user engagement deteriorate due to limited input variation
Solution Approach 1:
The patent implements multi-functionality by integrating sound recognition capabilities into the existing AR system. The sound input device and sound recognition means can recognize various types of sounds (voice commands, clapping, whistling) and translate them into different input operations, enabling diverse interactiveness without requiring separate devices for each input method. This universal approach significantly increases input method variation while adding moderate system complexity.
Data Source
AI summary
An image taken by a real camera is repeatedly obtained, and position and orientation information determined in accordance with a position and an orientation of a real camera in a real space is repeatedly calculated. A virtual object or a letter to be additionally displayed on the taken image is set as an additional display object, and based on a result of recognition of a sound inputted into a sound input device, at least one selected from the group consisting of a display position, an orientation, and a display form of the additional display object is set. A combined image repeatedly generated by superimposing on the taken image the set additional display object with reference to a position in the taken image in accordance with the position and orientation information is displayed on a display device.


