Facial Information Extraction Using Voice Keyword Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice recognition systems fail to accurately acquire facial information from images containing multiple facial images, as they cannot determine specific facial information based on voice queries.

Innovation Solution

A method and apparatus that acquire to-be-processed voice information and images, perform voice recognition to generate query information, semantically recognize the query to create a keyword set, and use this set to process the image for accurate facial information extraction, including position information and personal details.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice recognition is used to query facial information from images with multiple faces, then the ease of operation is improved, but the measurement precision deteriorates because the system cannot accurately identify which specific facial image the user wants

Engineering Contradiction:
Improveease of operationVSAvoidmeasurement precision
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent segments the image processing task into multiple stages: first detecting all facial images in the input image, then for each facial image extracting feature vectors and matching them with voice-recognized keyword vectors. This segmentation allows the system to handle multiple faces systematically and accurately identify the target face based on voice query, resolving the contradiction between ease of voice operation and precise face identification.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces feature vectors as an intermediary representation between the visual facial images and the voice-recognized keywords. By converting both facial images and voice queries into comparable vector representations in the same feature space, the system enables accurate matching and identification, solving the precision problem while maintaining voice-based ease of operation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If semantic recognition and keyword set generation are performed, then the measurement precision of facial information acquisition is improved, but the device complexity increases due to multiple processing stages

Engineering Contradiction:
Improvemeasurement precisionVSAvoiddevice complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent designs a unified processing framework where the same feature extraction and vector matching mechanism serves multiple functions: detecting facial images, extracting their features, matching with voice keywords, and identifying the target face. This multi-functionality reduces the need for separate specialized modules, thereby managing device complexity while maintaining high measurement precision.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent performs preliminary actions by pre-extracting feature vectors for all facial images in the input image before receiving the voice query. This preliminary processing prepares the data in advance, allowing the system to quickly and accurately match the voice-recognized keywords with the appropriate facial images, thus achieving high precision without requiring complex real-time processing during the query phase.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS10755079B2Method and apparatus for acquiring facial information
Publication Date: 2020.08.25 BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
  • US10755079B2 patent drawing
  • US10755079B2 patent drawing
  • US10755079B2 patent drawing

AI summary

Embodiments of the present disclosure disclose a method and apparatus for acquiring facial information. A specific embodiment of the method comprises: acquiring to-be-processed voice information and a to-be-processed image, and performing voice recognition on the to-be-processed voice information to acquire query information, the to-be-processed image comprising a plurality of facial images, and the query information being used to instruct querying facial information of a specified facial image in the to-be-processed image; recognizing semantically the query information to acquire a keyword set for querying the facial information; processing the to-be-processed image to acquire facial information in the to-be-processed image; and acquiring facial information corresponding to the to-be-processed voice information from the facial information by using the keyword set. The present embodiment realizes recognition on a plurality of facial images contained in a to-be-processed image and determines facial information by using a keyword, thereby improving the accuracy of acquiring a facial image by means of voice.