Mobile AR Navigation via Visual Recognition and Speech
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current mobile devices rely on manual input of locations for navigation, which can be cumbersome and inefficient, especially in situations where visual recognition of surroundings could provide quicker and more accurate location determination.
Innovation Solution
Equipping mobile devices with an electronic image sensor that captures images and overlays directional pointers on the screen, allowing users to identify locations and receive turn-by-turn directions based on recognized reference objects, such as buildings or signs, through image analysis and speech recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If manual input of location is used, then device complexity is reduced, but navigation efficiency and accuracy deteriorate
Solution Approach 1:
The mobile device is enhanced with multiple functional components including electronic image sensor, GPS receiver, access point communicator, and speech recognizer, allowing it to perform both visual recognition-based location determination and traditional manual input methods, thereby improving navigation efficiency while maintaining operational simplicity
Solution Approach 2:
The system introduces an image processing server as an intermediary that receives images from mobile devices, performs complex image analysis to identify reference objects and determine locations, and returns results to the mobile device, thereby enabling advanced visual recognition capabilities without significantly increasing the complexity of the mobile device itself
2Measurement precision
If visual recognition is implemented, then location determination accuracy is improved, but processing time increases
Solution Approach 1:
The system pre-loads and stores image data of reference objects (buildings, signs, landmarks) in databases at the image processing server. When a user captures an image, the system performs rapid matching against pre-stored reference images, significantly reducing processing time while maintaining high location determination accuracy
Solution Approach 2:
The image processing is divided into multiple stages: initial image capture and preprocessing on the mobile device, followed by transmission to the server for advanced analysis, and final result integration with GPS and access point data. This segmentation allows computationally intensive processing to occur on the server while keeping the mobile device responsive
3Adaptability or versatility
If electronic image sensor is added, then location identification capability is improved, but device complexity increases
Solution Approach 1:
The electronic image sensor is integrated into the existing mobile device architecture, allowing the same device to perform both traditional functions (calling, messaging, GPS navigation) and new visual recognition functions for location identification, thereby improving adaptability without requiring separate dedicated devices
Solution Approach 2:
The image processing server acts as an external intermediary that handles the complex computational tasks of image analysis and reference object matching. The mobile device with the added image sensor only needs to capture images and communicate with the server, significantly reducing the processing burden and effective complexity of the mobile device while still enabling advanced location identification capabilities
Data Source
AI summary
In one or more embodiments, one or more methods and/or systems described can perform producing a lattice of object hypotheses based on multiple reference objects from image information; receiving input speech information that includes a request for information associated with at least one reference object of the multiple reference objects; producing a lattice of speech hypotheses based on at least a first possible description included in the speech information; producing a lattice of scored semantic hypotheses based on at least the lattice of object hypotheses and the lattice of speech hypotheses; determining that a single semantic interpretation score of the lattice of scored semantic hypotheses exceeds a predetermined value; and providing requested information associated with the at least the first reference object of the plurality of reference objects.


