Portrait Image Recognition Using Voice and Focus Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional smart refrigerators with large screens face inefficiencies in recognizing multiple portraits on a home page, leading to increased recognition errors and prolonged processing times.

Innovation Solution

An image recognition method that determines the number of object boxes on a screen, sets a suitable time period, obtains external voice information and monitors focus events to select target object boxes, and performs deduplication processing to improve recognition efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If the smart refrigerator recognizes all portraits on the home page, then the recognition completeness is improved, but the recognition time increases and recognition errors increase

Engineering Contradiction:
Improverecognition completenessVSAvoidrecognition time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system performs preliminary actions by monitoring focus events and obtaining external voice information before actual recognition. It identifies and selects target object boxes that the user is interested in, preparing the recognition process in advance by filtering out only the necessary portraits rather than processing all portraits equally.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system extracts and selects specific target object boxes from the set of all object boxes based on user interaction (focus events) and voice commands. This extraction principle allows the system to isolate only the relevant portraits for recognition, eliminating unnecessary processing of other portraits.

Inventive Principle:
Principle #2Taking out (Extraction)

2Reliability

If the smart refrigerator recognizes all portraits on the home page, then the recognition completeness is improved, but the processing efficiency deteriorates

Engineering Contradiction:
Improverecognition completenessVSAvoidprocessing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The system performs preliminary filtering by monitoring focus events and obtaining voice information before recognition. This preliminary action identifies target object boxes in advance, so that the actual recognition process only needs to process the selected targets rather than all portraits, thereby improving processing efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system extracts target object boxes from the complete set of object boxes based on user focus events and voice commands. This extraction creates a reduced subset of portraits requiring recognition, maintaining completeness for user-interesting portraits while dramatically reducing overall processing workload.

Inventive Principle:
Principle #2Taking out (Extraction)

3Reliability

If the smart refrigerator processes more object boxes, then the recognition coverage is improved, but the recognition accuracy deteriorates due to increased errors

Engineering Contradiction:
Improverecognition coverageVSAvoidrecognition accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The system extracts and processes only the target object boxes that are relevant to user interest, as determined by focus events and voice information. This selective extraction maintains recognition coverage for important portraits while reducing the total number of processing operations, thereby improving recognition accuracy by avoiding errors from processing excessive non-target portraits.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12475700B2Image recognition
Publication Date: 2025.11.18 SHENZHEN TCL NEW-TECH CO LTD
  • US12475700B2 patent drawing
  • US12475700B2 patent drawing

AI summary

In an image recognition method, a unit duration is set according to the actual number of object boxes. External voice information is obtained or a focus event is monitored within the unit duration. One or more target object boxes are selected according to at least one of the external voice information or the focus event. Deduplication processing is performed on target object images respectively contained in the target object boxes.