Speaker recognition accuracy

By fragmenting and embedding audio samples, combined with a neural network model, the problem of inaccurate speaker recognition in existing technologies has been solved, achieving higher recognition accuracy and device security.

CN116508097BActive Publication Date: 2026-05-26GOOGLE LLC

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GOOGLE LLC
Filing Date
2021-10-13
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies struggle to accurately identify speakers using limited audio data in voice-based user authentication, leading to insufficient accuracy in identity verification and potentially privacy and security issues.

Method used

By dividing audio samples into multiple segments, a set of candidate acoustic embeddings is generated, and embeddings that do not meet the criteria are removed to generate aggregated acoustic embeddings. Speakers are identified by matching distance thresholds and aggregated acoustic embeddings. The accuracy of recognition is improved by combining neural network acoustic models and spectrogram enhancement techniques.

Benefits of technology

It improves the accuracy of speaker recognition, reduces the possibility of false recognition, and enhances the security and privacy protection of the device.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116508097B_ABST
    Figure CN116508097B_ABST
Patent Text Reader

Abstract

A method (300) for generating accurate speaker representations for audio samples (202) includes receiving a first audio sample from a first speaker (10) and a second audio sample from a second speaker. The method includes dividing the corresponding audio samples into multiple audio segments (214). The method also includes generating a set of candidate acoustic embeddings (232) based on the multiple segments, wherein each candidate acoustic embedding includes a vector representation of acoustic features. The method further includes removing a subset of the candidate acoustic embeddings from the set of candidate acoustic embeddings. The method further includes generating an aggregated acoustic embedding (234) from the remaining candidate acoustic embeddings in the set of candidate acoustic embeddings after removing the subset.
Need to check novelty before this filing date? Find Prior Art