Portrait Image Recognition Using Voice and Focus Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional smart refrigerators with large screens face inefficiencies in recognizing multiple portraits on a home page, leading to increased recognition errors and prolonged processing times.
Innovation Solution
An image recognition method that determines the number of object boxes on a screen, sets a suitable time period, obtains external voice information and monitors focus events to select target object boxes, and performs deduplication processing to improve recognition efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If the smart refrigerator recognizes all portraits on the home page, then the recognition completeness is improved, but the recognition time increases and recognition errors increase
Solution Approach 1:
The system performs preliminary actions by monitoring focus events and obtaining external voice information before actual recognition. It identifies and selects target object boxes that the user is interested in, preparing the recognition process in advance by filtering out only the necessary portraits rather than processing all portraits equally.
Solution Approach 2:
The system extracts and selects specific target object boxes from the set of all object boxes based on user interaction (focus events) and voice commands. This extraction principle allows the system to isolate only the relevant portraits for recognition, eliminating unnecessary processing of other portraits.
2Reliability
If the smart refrigerator recognizes all portraits on the home page, then the recognition completeness is improved, but the processing efficiency deteriorates
Solution Approach 1:
The system performs preliminary filtering by monitoring focus events and obtaining voice information before recognition. This preliminary action identifies target object boxes in advance, so that the actual recognition process only needs to process the selected targets rather than all portraits, thereby improving processing efficiency.
Solution Approach 2:
The system extracts target object boxes from the complete set of object boxes based on user focus events and voice commands. This extraction creates a reduced subset of portraits requiring recognition, maintaining completeness for user-interesting portraits while dramatically reducing overall processing workload.
3Reliability
If the smart refrigerator processes more object boxes, then the recognition coverage is improved, but the recognition accuracy deteriorates due to increased errors
Solution Approach 1:
The system extracts and processes only the target object boxes that are relevant to user interest, as determined by focus events and voice information. This selective extraction maintains recognition coverage for important portraits while reducing the total number of processing operations, thereby improving recognition accuracy by avoiding errors from processing excessive non-target portraits.
Data Source
AI summary
In an image recognition method, a unit duration is set according to the actual number of object boxes. External voice information is obtained or a focus event is monitored within the unit duration. One or more target object boxes are selected according to at least one of the external voice information or the focus event. Deduplication processing is performed on target object images respectively contained in the target object boxes.

