Machine Learning Audio Quality Classification for Video Game Dialogue

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Video games often require vast amounts of dialogue recordings, leading to quality assurance challenges due to varying audio quality from different recording studios, resulting in significant developer time spent on cleanup and quality assessment.

Innovation Solution

A data processing apparatus utilizing machine learning models to classify audio data quality, storing identification data for low-quality recordings for further analysis or replacement, and providing a system for efficient quality assurance and enhancement.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If manual quality assurance processes are used to assess and clean up dialogue recordings, then audio quality can be ensured, but significant developer time is spent on quality assurance and cleanup

Engineering Contradiction:
Improveaudio qualityVSAvoiddeveloper time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system enables self-service quality assurance by training machine learning models to automatically assess and classify dialogue recording quality. The models analyze audio properties such as noise levels, clarity, and consistency, then prioritize recordings for cleanup without human intervention, allowing the system to serve its own quality control needs

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces manual mechanical quality assessment processes with automated machine learning models. These models substitute human developers' time-consuming manual review and cleanup work with algorithmic analysis of audio properties, automatically identifying low-quality recordings that need attention

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Quantity of substance

If vast amounts of dialogue recordings are recorded for video games, then comprehensive dialogue coverage is achieved, but quality assurance becomes increasingly difficult and time-consuming

Engineering Contradiction:
Improvedialogue recording volumeVSAvoidquality assurance complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system segments the large volume of dialogue recordings into manageable groups based on quality classification. Machine learning models analyze audio properties and categorize recordings by quality level, creating segmented lists that prioritize which recordings need cleanup attention, making the overwhelming task of reviewing thousands of recordings tractable

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes the parameters used for quality assessment by training machine learning models to evaluate specific audio properties such as noise floor levels, signal-to-noise ratio, spectral characteristics, and temporal consistency. These parameter-based classifications enable systematic handling of large dialogue volumes

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If dialogue recordings are made at different recording studios with different hardware and sound proofing, then recording flexibility is improved, but audio quality consistency deteriorates

Engineering Contradiction:
Improverecording location flexibilityVSAvoidaudio quality consistency
Core Design Contradiction:
Adaptability or versatilityVSStability of the object's composition

Solution Approach 1:

The machine learning models act as an intermediary between diverse recording environments and quality standards. The models analyze audio properties from recordings made at different studios with varying hardware and acoustics, identifying consistent quality patterns across environments and prioritizing recordings that deviate from expected quality thresholds

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system adjusts quality assessment parameters to account for different recording environments by training models to recognize studio-specific acoustic characteristics. This allows the system to maintain quality consistency standards across diverse locations while adapting to the unique properties of each recording environment

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12070688B2Apparatus and method for audio data analysis
Publication Date: 2024.08.27 SONY INTERACTIVE ENTERTAINMENT LLC
  • US12070688B2 patent drawing
  • US12070688B2 patent drawing
  • US12070688B2 patent drawing

AI summary

A data processing apparatus includes input circuitry to receive audio data for a plurality of respective dialogue recordings for a video game, classification circuitry comprising one or more machine learning models to receive at least a portion of the audio data for each dialogue recording and trained to output classification data indicative of a quality classification of a dialogue recording in dependence upon one or more properties of the audio data for the dialogue recording, and storage circuitry to store identification data for one or more of the plurality of dialogue recordings in dependence upon the classification data.