Speaker Recognition Registration with Synthesized Noise

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speaker recognition systems face performance degradation due to mismatches between registration and test environments, particularly when exposed to different noise types and levels, leading to recognition errors.

Innovation Solution

A method that synthesizes a speech signal with a preset noise signal during registration, generating feature vectors that are robust against various noise types, using techniques like convolutional neural networks and domain transformations to improve verification performance across different environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speaker recognition is performed using speech signals recorded in a controlled registration environment, then the recognition accuracy is high under ideal conditions, but the performance degrades significantly when tested in noisy environments

Engineering Contradiction:
Improverecognition accuracyVSAvoidenvironmental adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies preliminary action by synthesizing noise signals and mixing them with speech signals during the registration phase. This prepares the speaker model in advance to account for various noise conditions that may occur during testing, thereby improving environmental adaptability without sacrificing recognition accuracy. The noise synthesis and signal mixing are performed beforehand to create robust speaker representations.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the parameters of the speech signal by introducing synthesized noise components with varying noise levels and types during registration. This transforms the registration process from using clean speech signals to using noise-contaminated signals, enabling the system to adapt to different acoustic environments while maintaining recognition performance.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If noise signals are synthesized and mixed with speech signals during registration, then robustness against noise is improved, but the complexity of the registration process increases

Engineering Contradiction:
Improvenoise robustnessVSAvoidregistration process complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary noise synthesis component that generates artificial noise signals to be mixed with speech signals during registration. This intermediary element enables the system to simulate various acoustic environments without requiring actual recordings from noisy environments, thereby improving noise robustness while keeping the system architecture manageable.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates copies of speech signals with added synthesized noise components during registration. Instead of requiring multiple actual recordings from different noisy environments, the system generates synthetic copies that mimic various noise conditions, simplifying the registration process while improving robustness.

Inventive Principle:
Principle #26Copying

Data Source

PatentEP3706117B1Method with speaker recognition registration and corresponding non-transitory computer-readable storage medium
Publication Date: 2022.05.11 SAMSUNG ELECTRONICS CO LTD
  • EP3706117B1 patent drawingFigure 1
  • EP3706117B1 patent drawingFigure 2
  • EP3706117B1 patent drawingFigure 3

AI summary

A method with registration includes: receiving a speech signal of a speaker; synthesizing the received speech signal and a noise signal to generate a synthesized signal; generating a feature vector based on the synthesized signal; and constructing a registration database (DB) corresponding to the speaker based on the generated feature vector.