Image Object Recognition with Natural-Language Query Disambiguation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models for object recognition from images require repetitive question-answering interactions with users due to ambiguous natural language queries, leading to inefficiencies and complexity in the recognition process.

Innovation Solution

A method and apparatus that utilize a machine learning model to recognize objects from images based on natural language queries, determining if the number of instances satisfies a predetermined condition, and if not, engaging in additional question-answering processes to refine the query, with a self-evolving training mechanism to improve accuracy and reduce interaction complexity.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If the machine learning model uses natural language queries for object recognition, then the ease of operation is improved, but the device complexity increases due to ambiguous queries requiring repetitive confirmation

Engineering Contradiction:
Improveease of operationVSAvoiddevice complexity
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system performs preliminary analysis of the natural language query to determine whether it is ambiguous before initiating the recognition process. This preliminary action allows the system to prepare appropriate disambiguation strategies in advance, reducing the need for repetitive back-and-forth interactions with the user and simplifying the overall operational flow.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system implements a feedback mechanism where the machine learning model provides recognition results to the user, and based on user confirmation or correction, iteratively refines the recognition process. This feedback loop resolves ambiguity efficiently by involving the user only when necessary, rather than requiring repetitive confirmation for every query.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If the machine learning model repeatedly confirms recognition targets with the user, then the measurement precision is improved, but the loss of time increases due to multiple interactions

Engineering Contradiction:
Improvemeasurement precisionVSAvoidloss of time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system employs self-service mechanisms where the machine learning model autonomously resolves ambiguous queries by analyzing contextual information, generating multiple candidate interpretations, and selecting the most probable one without requiring user intervention. User confirmation is requested only when the system's confidence threshold is not met, significantly reducing interaction time while maintaining precision.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system dynamically adjusts the confidence threshold parameter based on the ambiguity level of the query. For highly ambiguous queries, a lower threshold allows faster automatic resolution, while for critical recognitions, a higher threshold ensures greater precision by triggering user confirmation only when absolutely necessary.

Inventive Principle:
Principle #35Parameter changes

3Manufacturing precision

If the machine learning model improves understanding ability to process images accurately, then the manufacturing precision is improved, but the device complexity increases due to additional training mechanisms

Engineering Contradiction:
Improvemanufacturing precisionVSAvoiddevice complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The system merges the object recognition model with a natural language processing model into a unified architecture that jointly processes both image and query data. This integration shares computational resources and parameters between the two functions, improving recognition accuracy while avoiding the overhead of completely separate systems.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The machine learning model is designed with multi-functionality to handle both standard object recognition and ambiguous query disambiguation tasks using the same underlying architecture. This universal model reduces device complexity by eliminating the need for specialized components for different recognition scenarios.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250259422A1Method, apparatus, device and medium for object recognition from image
Publication Date: 2025.08.14 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US20250259422A1 patent drawing
  • US20250259422A1 patent drawing
  • US20250259422A1 patent drawing

AI summary

Methods, apparatuses, devices and media for object recognition from an image are provided. In a method, a query expressed in a natural language is received, the query specifying an attribute of an object to be recognized from the image. At least one instance of the object is recognized from the image based on the query using a machine learning model. The at least one instance is provided in response to determining that a number of the at least one instance satisfies a predetermined condition. By the example implementations of the subject matter described herein, rounds of conversation between the machine learning model and a user may be significantly reduced, thereby the object is recognized from the image in a simpler and more efficient way.