Voice Model Transfer Across Devices

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice recognition systems require users to train each device individually, which is time-consuming and inefficient, especially when environmental conditions vary across different systems.

Innovation Solution

A method that estimates environment-specific alterations in user sounds and uses a user-dependent audio model stored in a multi-system store to identify users across multiple systems, eliminating the need for repeated training by compensating for environmental changes in the sound before comparison.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If a user trains each voice recognition system individually, then the system can accurately identify the user's voice, but the user must invest considerable time in training each system

Engineering Contradiction:
Improvevoice identification accuracyVSAvoiduser training time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary voice training action on a first voice recognition system, then transfers the trained voice model to a second system. This preliminary action eliminates the need for repeated training on each system, reducing user time investment while maintaining identification accuracy across multiple devices

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a copy of the trained voice model from the first voice recognition system and applies it to the second system. This copying mechanism allows the user's voice characteristics to be recognized on multiple systems without retraining, resolving the contradiction between accuracy and training time

Inventive Principle:
Principle #26Copying

2Adaptability or versatility

If the training process is customized for each system, then the system can adapt to system-specific characteristics, but the training process becomes different for each system requiring more user effort

Engineering Contradiction:
Improvesystem-specific adaptationVSAvoidtraining process simplicity
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent makes the trained voice model universal by transferring it from a first voice recognition system to a second system. The model adapts to different systems while maintaining consistent training procedures, eliminating the need for system-specific training processes and improving ease of operation

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Adaptability or versatility

If voice recognition is implemented across multiple systems, then user identification can be achieved on any device, but each system would traditionally require separate training

Engineering Contradiction:
Improvemulti-system compatibilityVSAvoidtraining management complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent copies the trained voice model to multiple systems, enabling multi-system compatibility without requiring separate training for each device. This copying approach simplifies training management while maintaining the ability to identify users across different devices

Inventive Principle:
Principle #26Copying

Data Source

PatentEP3248189B1Environment adjusted speaker identification
Publication Date: 2023.05.03 MICROSOFT TECHNOLOGY LICENSING LLC
  • EP3248189B1 patent drawingFigure 1
  • EP3248189B1 patent drawingFigure 2
  • EP3248189B1 patent drawingFigure 3~4

AI summary

Computerized estimation of an identity of a user of a computing system. The system estimates environment-specific alterations of a received user sound that is received at the computing system. The system estimates whether the received user sounds is from a particular user by use of a corresponding user-dependent audio model. The user-dependent audio model may be stored in a multi-system store accessible such that the method may be performed for a given user across multiple systems and on a system that the user has never before trained to recognize the user. This reduces or even eliminates the need for a user to train a system to recognize the voice of a user, and allows multiple systems to take advantage of previous training performed by the user.