Automated Pipeline Selection for Audio Asset Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Interactive software applications, such as video games, face challenges in generating high-quality synthesized audio streams using text-to-speech (TTS) and voice conversion (VC) techniques, as existing methods lack efficient automated pipeline selection based on audio stream features, leading to suboptimal audio asset quality.

Innovation Solution

The implementation of an automated pipeline selection system that analyzes audio stream features like size, sampling rate, pitch, and speaker characteristics to select the most suitable TTS or VC pipeline for training, utilizing neural networks and rule engines to determine the best pipeline configuration for generating high-quality audio assets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If automated pipeline selection is implemented based on audio stream features, then audio asset quality is improved, but system complexity increases

Engineering Contradiction:
Improveaudio asset qualityVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent introduces an automated pipeline selection system that acts as an intermediary between audio stream input and TTS/VC pipeline processing. This selection system analyzes audio stream features (size, sampling rate, pitch, speaker characteristics) and automatically determines the most suitable pipeline configuration, thereby improving audio asset quality without requiring manual intervention while managing system complexity through automation.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If multiple TTS and VC pipelines are maintained for different audio conditions, then adaptability is improved, but device complexity increases

Engineering Contradiction:
Improvepipeline adaptabilityVSAvoidpipeline configuration complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements a dynamic pipeline selection mechanism that automatically adapts the TTS/VC processing configuration based on real-time analysis of audio stream features. Instead of maintaining fixed, separate pipelines for different conditions, the system dynamically selects and configures the appropriate pipeline based on the specific characteristics of each audio stream, thereby achieving high adaptability while simplifying the overall system architecture.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes operational parameters (pipeline selection, configuration settings) based on audio stream features such as size, sampling rate, pitch, and speaker characteristics. By dynamically adjusting pipeline parameters rather than maintaining multiple fixed pipelines, the system achieves versatility while managing complexity through parameter-based adaptation.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If manual pipeline selection is used for TTS and VC processing, then ease of operation is maintained, but productivity decreases

Engineering Contradiction:
Improveaudio asset generation efficiencyVSAvoidpipeline selection complexity
Core Design Contradiction:
ProductivityVSEase of operation

Solution Approach 1:

The patent implements a self-service automated pipeline selection system that independently analyzes audio stream features and selects the appropriate TTS/VC pipeline without requiring manual intervention. The system extracts relevant features (size, sampling rate, pitch, speaker characteristics) and automatically determines the optimal pipeline configuration, thereby significantly improving productivity while maintaining ease of operation through complete automation of the selection process.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS12159618B2Automated pipeline selection for synthesis of audio assets
Publication Date: 2024.12.03 ELECTRONIC ARTS INC
  • US12159618B2 patent drawing
  • US12159618B2 patent drawing
  • US12159618B2 patent drawing

AI summary

An example method of automated selection of audio asset synthesizing pipelines includes: receiving an audio stream comprising human speech; determining one or more features of the audio stream; selecting, based on the one or more features of the audio stream, an audio asset synthesizing pipeline; training, using the audio stream, one or more audio asset synthesizing models implementing respective stages of the selected audio asset synthesizing pipeline; and responsive to determining that a quality metric of the audio asset synthesizing pipeline satisfies a predetermined quality condition, synthesizing one or more audio assets by the selected audio asset synthesizing pipeline.