Natural-Language Query Feedback for Ambiguous Image Object Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models for object recognition from images require repetitive question-answering interactions with users due to ambiguous natural language queries, leading to inefficiency and complexity in the recognition process.

Innovation Solution

A method and apparatus that utilize a machine learning model to recognize objects based on a natural language query, determining if the number of instances satisfies a predetermined condition, and if not, engaging in additional question-answering processes to refine the recognition, with a self-evolving training mechanism to improve accuracy and reduce interactions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a machine learning model uses natural language queries for object recognition, then the system becomes more user-friendly and flexible, but the model requires repetitive question-answering interactions due to ambiguous queries

Engineering Contradiction:
Improvenatural language query processingVSAvoidquestion-answering interaction process
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system performs preliminary analysis of the natural language query to determine whether it is ambiguous before initiating the recognition process. By evaluating query clarity in advance, the system can prepare appropriate responses or clarification requests, reducing the need for multiple interactive rounds and streamlining the overall process.

Inventive Principle:
Principle #10Preliminary action

2Productivity

If the machine learning model processes ambiguous queries without clarification, then the recognition process is faster, but the accuracy of object recognition decreases

Engineering Contradiction:
Improverecognition process speedVSAvoidobject recognition accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The system implements a feedback mechanism where the model evaluates the clarity of the natural language query and provides feedback on whether clarification is needed. This feedback loop allows the system to maintain high accuracy by identifying ambiguous queries that require additional information, while avoiding unnecessary interactions when queries are already clear.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If the system requires multiple rounds of question-answering to clarify ambiguous queries, then the recognition accuracy improves, but the time required for the process increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidquestion-answering interaction time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system employs self-service mechanisms where the machine learning model automatically evaluates query ambiguity and determines whether clarification is needed without requiring manual intervention or multiple interactive rounds. This self-assessment capability reduces time loss by eliminating unnecessary question-answering cycles while maintaining recognition accuracy.

Inventive Principle:
Principle #25Self-service

Data Source

PatentEP4600842A1Method, apparatus, device and medium for object recognition from image
Publication Date: 2025.08.13 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • EP4600842A1 patent drawingFigure 1~2
  • EP4600842A1 patent drawingFigure 3~4
  • EP4600842A1 patent drawingFigure 5~6

AI summary

Methods, apparatuses, devices and media for object recognition from an image are provided. In a method, a query expressed in a natural language is received, the query specifying an attribute of an object to be recognized from the image. At least one instance of the object is recognized from the image based on the query using a machine learning model. The at least one instance is provided in response to determining that a number of the at least one instance satisfies a predetermined condition. By the example implementations of the subject matter described herein, rounds of conversation between the machine learning model and a user may be significantly reduced, thereby the object is recognized from the image in a simpler and more efficient way.