Object Recognition Using Voiceprint and Position Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In voiceprint recognition systems, distinguishing speakers with similar voiceprint features is challenging, leading to reduced accuracy in identifying speakers, especially in multi-speaker environments.

Innovation Solution

An object recognition method that combines speech information and position information using a trained voiceprint matching model to extract voiceprint features and calculate a confidence value, allowing for more accurate identification of target objects by integrating voice confidence values with position information.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If voiceprint matching is used to identify speakers, then speaker identification can be performed, but accuracy decreases when voiceprint features are similar

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoidrecognition reliability in similar voiceprint scenarios
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent combines voiceprint feature matching with position information to form a composite recognition system. The voiceprint module extracts acoustic features while the position module determines spatial location, and both results are integrated through a scoring mechanism to achieve more reliable speaker identification, especially when voiceprint features are similar.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces a score calculation module as an intermediary that processes both voiceprint matching results and position information. This intermediary computes a comprehensive recognition score by weighting and combining multiple factors, thereby mediating between the voiceprint features and final identification decisions to improve accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Device complexity

If only voiceprint features are used for recognition, then the system is simple, but it cannot distinguish speakers with similar voiceprint patterns

Engineering Contradiction:
Improverecognition system complexityVSAvoidspeaker distinction accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent adds a spatial dimension to the voiceprint recognition system by incorporating position information. Instead of relying solely on acoustic feature space, the system now operates in a combined space of voiceprint characteristics and spatial location, enabling differentiation of speakers with similar voiceprints through their distinct positional information.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Productivity

If voiceprint matching is performed in multi-speaker environments, then speaker identification is possible, but accuracy reduces due to similar voiceprint features

Engineering Contradiction:
Improvespeaker identification capabilityVSAvoidrecognition accuracy in complex environments
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent segments the speaker identification process into distinct functional modules: voiceprint feature extraction, position information acquisition, score calculation, and final recognition decision. This segmentation allows each module to specialize in specific tasks and enables the system to handle multi-speaker environments more effectively by independently processing and integrating different information sources.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS11289072B2Object recognition method, computer device, and computer-readable storage medium
Publication Date: 2022.03.29 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US11289072B2 patent drawing
  • US11289072B2 patent drawing
  • US11289072B2 patent drawing

AI summary

An object recognition method is provided. The method includes obtaining speech information of a target object in a current speech environment and position information of the target object; extracting voiceprint feature from the speech information based on a trained voiceprint matching model, to obtain voiceprint feature information; obtaining a voice confidence value corresponding to the voiceprint feature information; and obtaining an object recognition result of the target object based on the voice confidence value, the position information, and the voiceprint feature information.