Ambient Voice Capture for In-Game Character Synthesis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video game technologies lack immersive experiences due to limited utilization of ambient audio data from environments, resulting in standard and unpersonalized non-player character (NPC) voices.

Innovation Solution

A computer system that uses ambient noise from microphones to distinguish user and background voices, derive voice fingerprints, and input these parameters into a generative AI model for synthesizing speech that mimics the background voices, enhancing realism and personalization in video games.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If ambient noise is captured and processed to create personalized NPC voices, then voice quality and immersion are improved, but system complexity and processing requirements increase

Engineering Contradiction:
Improvevoice qualityVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary voice fingerprint extraction and phoneme isolation from ambient noise during gameplay, storing these parameters for later speech synthesis. This preliminary processing enables quick generation of personalized NPC voices without real-time complexity

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary generative AI model that bridges the gap between captured ambient voice parameters and synthesized speech output. This model processes the complex transformation from raw audio parameters to natural-sounding speech, isolating the complexity from the main game engine

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If phonemes are isolated from background voices in real-time, then speech synthesis accuracy is improved, but processing time and computational resources increase

Engineering Contradiction:
Improvespeech synthesis accuracyVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system extracts and stores phoneme parameters from background voices in advance during gameplay sessions. These pre-processed phoneme representations are saved for rapid retrieval and synthesis, eliminating the need for real-time phoneme isolation during critical game moments

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies partial phoneme isolation by focusing on extracting only the essential phoneme parameters needed for speech synthesis rather than complete voice replication. This selective processing reduces computational overhead while maintaining sufficient accuracy for NPC dialogue

Inventive Principle:
Principle #16Partial or excessive action

3Adaptability or versatility

If multiple background voices are distinguished and processed, then voice diversity and realism are improved, but algorithm complexity increases

Engineering Contradiction:
Improvevoice diversityVSAvoidalgorithm complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system segments the audio environment into distinct voice sources using voice fingerprinting technology. Each background voice is identified and processed separately, allowing the system to handle multiple voices independently and combine them appropriately for diverse NPC dialogue

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms complex voice distinction problems into parameter-based processing by extracting key acoustic features and phoneme representations. This parameter transformation simplifies the comparison and differentiation of multiple background voices, reducing algorithmic complexity while maintaining voice diversity

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240304174A1Ambient Noise Capture for Speech Synthesis of In-Game Character Voices
Publication Date: 2024.09.12 SONY INTERACTIVE ENTERTAINMENT LLC
  • US20240304174A1 patent drawing
  • US20240304174A1 patent drawing
  • US20240304174A1 patent drawing

AI summary

Methods and systems are presented for discerning background voices from ambient noise in a video game player's surroundings, isolating the voices using distinct phonemes within the voices, and running the phonemes through a generative artificial intelligence (AI) model to produce synthesized speech in the video game. The synthesized speech can be spoken by players or non-player characters, modified in age, gender, and other attributes of the avatar renderings. The text for the synthesized speech can be for static scripts or alterations based on gameplay. The player can select from different background voices and apply them to in-game characters and elements as desired.