Spatial Character Voice Processing for Private Game Audio

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing game scenes provide a simple and direct voice audio processing mode that lacks stereoscopic spatial representation and privacy, resulting in flat sound effects and exposure of real user identities through similar voice timbres.

Innovation Solution

An audio processing method that converts voice audio to match character attributes of virtual objects, incorporating spatial position information for enhanced realism and privacy, using terminals and servers to process and transmit target audio and position data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If simple and direct voice audio processing mode is used, then processing complexity is reduced, but stereoscopic spatial representation and privacy are lost

Engineering Contradiction:
Improveprocessing complexityVSAvoidspatial representation and privacy
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The voice audio processing is segmented into multiple independent components: spatial position information extraction, voice timbre analysis, character attribute matching, and audio synthesis. Each component processes specific aspects separately before combining them into the final output, maintaining complexity manageability while achieving comprehensive processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A character attribute model serves as an intermediary between the user's voice and the virtual object's audio output. The model transforms real voice characteristics into fictional character attributes, preserving privacy while maintaining spatial representation. This intermediary layer decouples the direct connection between user identity and audio output.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of time

If voice audio is transmitted directly without conversion, then processing time is reduced, but user identity privacy is exposed

Engineering Contradiction:
Improveprocessing timeVSAvoididentity exposure
Core Design Contradiction:
Loss of timeVSObject-affected harmful factors

Solution Approach 1:

Character attribute profiles are pre-computed and stored before actual voice communication occurs. When voice audio needs processing, the system directly matches against pre-prepared character attributes rather than analyzing from scratch, significantly reducing processing time while maintaining privacy protection through the character attribute layer.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If voice timbre is converted to match character attributes, then privacy is protected and realism is enhanced, but processing complexity increases

Engineering Contradiction:
Improveprivacy protection and realismVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The system changes specific acoustic parameters of the voice audio (timbre, pitch, resonance characteristics) to match predefined character attribute parameters. By focusing transformation on key acoustic parameters rather than complete waveform manipulation, the system achieves realistic character voice representation while managing processing complexity through targeted parameter adjustment.

Inventive Principle:
Principle #35Parameter changes

4Loss of information

If spatial position information is incorporated into audio processing, then stereoscopic spatial sense is enhanced, but data transmission volume increases

Engineering Contradiction:
Improvespatial senseVSAvoiddata transmission volume
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

Spatial position information is extracted and transmitted as separate metadata alongside the audio data, rather than embedding it within the audio stream itself. This extraction approach allows efficient transmission of positional data using compact coordinate representations, minimizing the increase in overall data transmission volume while maintaining complete spatial information.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS12383832B2Audio processing method and apparatus
Publication Date: 2025.08.12 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US12383832B2 patent drawing
  • US12383832B2 patent drawing
  • US12383832B2 patent drawing

AI summary

An audio processing method and apparatus are provided. The method includes: obtaining a voice audio of a first game user and spatial position information of a first virtual object controlled by the first game user in a game scene; performing conversion processing on the voice audio of the first game user to obtain a target audio matching a character attribute of the first virtual object; and transmitting the target audio and the spatial position information of the first virtual object to a second terminal such that the second terminal plays the target audio according to the spatial position information of the first virtual object, a second virtual object controlled by a second game user using the second terminal and the first virtual object being in a same game scene.