Synthesizing Children Speech Models via Adult Transformation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition models for adults are ineffective for children's speech due to its unique characteristics, such as higher pitch and greater variability, which are not accounted for in existing adult models.

Innovation Solution

A transformation matrix derived from adult male and female speech models is modified and applied to produce a children's speech model, utilizing a fractional exponent to adjust the transformation matrix, allowing for the conversion of adult speech models to children's speech models effectively.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If adult speech models are used for children's speech recognition, then the model is already available and requires no additional training data, but the recognition accuracy deteriorates due to pitch and variability differences

Engineering Contradiction:
Improvemodel availabilityVSAvoidrecognition accuracy
Core Design Contradiction:
Ease of manufactureVSMeasurement precision

Solution Approach 1:

The patent applies parameter changes by modifying the acoustic model parameters through a transformation matrix derived from adult speech data. The transformation adjusts pitch and variability parameters to match children's speech characteristics, enabling accurate recognition without collecting new children's speech data. This resolves the contradiction by changing model parameters rather than model structure.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces an intermediary transformation matrix that bridges adult speech models and children's speech characteristics. This matrix, derived from comparing adult male and female speech patterns, acts as a mediator to adapt adult models for children without requiring direct children's speech training data, thus maintaining model availability while improving accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If a dedicated children's speech model is created, then recognition accuracy improves, but the complexity and resource requirements increase

Engineering Contradiction:
Improverecognition accuracyVSAvoidmodel complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent achieves universality by creating a single transformation matrix that can adapt adult speech models for children across different languages and contexts. This universal adaptation mechanism avoids the need for separate dedicated children's models for each language, reducing overall system complexity while maintaining high recognition accuracy.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Ease of operation

If the transformation matrix is applied without modification, then the conversion from adult to children's speech model is simple, but the recognition performance is insufficient

Engineering Contradiction:
Improvetransformation simplicityVSAvoidrecognition performance
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent modifies transformation parameters by adjusting the exponent applied to the transformation matrix. By optimizing this parameter (finding the optimal exponent value), the system achieves both computational simplicity and high recognition performance, resolving the contradiction between ease of operation and recognition precision.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP2332139B1Method for creating a speech model
Publication Date: 2015.10.21 ROSETTA STONE LTD
  • EP2332139B1 patent drawingFigure 1
  • EP2332139B1 patent drawingFigure 2
  • EP2332139B1 patent drawingFigure 3

AI summary

A transformation can be derived which would represent that processing required to convert a male speech model to a female speech model. That transformation is subjected to a predetermined modification, and the modified transformation is applied to a female speech model to produce a synthetic children's speech model. The male and female models can be expressed in terms of a vector representing key values defining each speech model and the derived transformation can be in the form of a matrix that would transform the vector of the male model to the vector of the female model. The modification to the derived matrix comprises applying an exponential p which has a value greater than zero and less than 1.