Voiceprint Recognition Model for Short Recordings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voiceprint recognition systems experience high recognition error rates for short recordings due to the limited ability of universal background models to capture subtle differences, leading to poor performance in short voice authentication scenarios.

Innovation Solution

An electronic device and method that involves framing processing of voice data, extraction of acoustic features using a predetermined filter, and pairwise coupling of these features with pre-stored units, followed by input into a pre-trained identity verification model for improved recognition, specifically utilizing a deep convolution neural network for enhanced performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a universal background model is used for voiceprint recognition, then recognition performance is good for long recordings, but recognition error rate increases for short recordings

Engineering Contradiction:
Improverecognition accuracyVSAvoidadaptability to different recording lengths
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent segments the voiceprint recognition model into multiple specialized components: a long recording recognition model and a short recording recognition model. Each model is trained specifically on its designated recording length, allowing the system to adapt to different recording durations without compromising accuracy. The segmentation enables the system to select the appropriate model based on the input recording length, thereby resolving the contradiction between maintaining high recognition accuracy and adapting to varying recording lengths.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent implements a dynamic model selection mechanism that adjusts the recognition approach based on the detected recording length. The system dynamically switches between the long recording recognition model and the short recording recognition model according to the actual input, making the system flexible and adaptive to different scenarios while maintaining optimal recognition performance for each type of recording.

Inventive Principle:
Principle #15Dynamics

2Device complexity

If a universal background model with limited parameters is used, then the model structure remains simple, but the ability to capture subtle differences in short recordings is insufficient

Engineering Contradiction:
Improvemodel structure complexityVSAvoidfeature extraction precision
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent divides the feature extraction and recognition process into separate specialized models for long and short recordings. Each model has its own optimized feature extraction capabilities tailored to the characteristics of its target recording length. This segmentation allows each model to develop specialized precision for capturing subtle differences in its designated domain without requiring a single overly complex universal model.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent creates separate copied model structures for long and short recording recognition, each independently trained and optimized. Rather than attempting to enhance a single universal model with increasing complexity, the system replicates the core recognition architecture multiple times with specialized training data, achieving high measurement precision through specialized copies rather than a single complex model.

Inventive Principle:
Principle #26Copying

Data Source

PatentEP3460793B1Electronic apparatus, identity verification method and system, and computer-readable storage medium
Publication Date: 2023.04.05 PING AN TECH (SHENZHEN) CO LTD
  • EP3460793B1 patent drawingFigure 1
  • EP3460793B1 patent drawingFigure 2
  • EP3460793B1 patent drawing

AI summary

The disclosure discloses an electronic device, a method and system of identity verification and a computer readable storage medium. The electronic device includes a memory and a processor; the system of identity verification is stored in the memory, and is executed by the processor to implement: after current voice data of a target user to be subjected to identity verification are received, carrying out framing processing on the current voice data according to preset framing parameters to obtain multiple voice frames; extracting preset types of acoustic features in all the voice frames by using a predetermined filter, and generating multiple observed feature units corresponding to the current voice data according to the extracted acoustic features; pairwise coupling all the observed feature units with pre-stored observed feature units respectively to obtain multiple groups of coupled observed feature units; inputting the multiple groups of coupled observed feature units into a preset type of identity verification model generated by pre-training to carry out the identity verification on the target user. The disclosure can reduce the error rate of short voice recognition.