Voice Query Enhancement via Focus Area Object Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current cognitive computing systems struggle to dynamically analyze the context of voice queries and correlate them with a user's focus area, leading to difficulties in interpreting incomplete or ambiguous queries due to factors like accents, speaking speeds, and non-verbal communication.

Innovation Solution

A system that analyzes voice queries using natural language processing, identifies objects in the user's focus area through sensors like cameras, determines confidence levels, and receives feedback to enhance query interpretation, generating accurate responses based on the identified objects and user input.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a virtual assistant expects a complete voice query from the user, then the accuracy of the response is improved, but the ease of operation deteriorates because users must formulate complete and precise queries

Engineering Contradiction:
Improvequery interpretation accuracyVSAvoiduser interaction convenience
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The system performs preliminary actions by capturing images of objects in the user's focus area before the user completes their voice query. This allows the system to pre-identify potential objects and prepare contextual information, so when the user speaks an incomplete query like 'What is this?', the system can immediately match it with the pre-captured object data and provide an accurate response without requiring the user to formulate a complete query.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces an intermediary mechanism - image capture and object recognition technology - that bridges the gap between incomplete user queries and accurate responses. The camera captures the object in the user's focus area, converts it into identifiable data, and uses this as an intermediary to link the ambiguous voice query with the specific object the user is referring to, thereby maintaining both ease of operation and response accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If the system captures an image of the object with a digital camera, then the ability to identify objects in the user's focus area is improved, but the device complexity increases

Engineering Contradiction:
Improveobject identification accuracyVSAvoidsystem structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system applies universality by integrating the camera into a multi-functional platform that serves both as a standard imaging device and as an object recognition component. The same camera hardware is used for various purposes including user authentication, environmental mapping, and object identification in focus areas. This eliminates the need for dedicated object recognition hardware, thereby improving object identification capability without proportionally increasing device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system merges the camera function with the voice processing system and object recognition algorithms into an integrated unit. Rather than adding a separate complex subsystem, the camera is combined with existing processing capabilities to create a unified object identification feature. This merging approach allows the system to leverage shared hardware and software resources, achieving improved object identification while minimizing the increase in overall system complexity.

Inventive Principle:
Principle #5Merging (Combining)

3Adaptability or versatility

If the system dynamically analyzes the context of voice queries and correlates them with the user's focus area, then the adaptability is improved, but the processing time increases

Engineering Contradiction:
Improvequery context understandingVSAvoidresponse time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The system performs preliminary context analysis by continuously capturing images and identifying objects in the user's focus area even before a voice query is spoken. This pre-processing of visual context allows the system to have object identification data ready in advance, so when the voice query arrives, the system can quickly correlate it with the pre-analyzed context rather than performing full analysis from scratch, thereby maintaining high adaptability while reducing response time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies local quality by focusing computational resources only on the specific region of interest - the user's focus area captured by the camera. Rather than analyzing the entire environment or performing exhaustive context analysis, the system concentrates processing power on identifying objects within the narrow field of view that the user is actually looking at. This localized approach enables rapid context understanding for the relevant area while minimizing overall processing time.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11481401B2Enhanced cognitive query construction
Publication Date: 2022.10.25 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11481401B2 patent drawing
  • US11481401B2 patent drawing
  • US11481401B2 patent drawing

AI summary

An embodiment for cognitively enhancing a search query is provided. The embodiment may include receiving a voice query from a user. The embodiment may also include analyzing the voice query. The embodiment may further include identifying an object within a focus area of the user based on the voice query. The embodiment may also include determining whether the identification of the object is confident, and in response to determining the identification of the object is not confident, receiving feedback from the user. In response to determining the identification of the object is confident, the embodiment may further include generating a relationship between a word in the voice query and the identified object. The embodiment may also include delivering an enhanced response to the user based on the identified object and the received feedback.