Multi-Source Audio Level Matching for Clearer Synthesized Voices

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing content distribution systems fail to adequately adjust the volume balance of multiple sound sources, leading to situations where the voice of one sound source overwhelms others, making it difficult to hear all sound sources clearly.

Innovation Solution

An information processing device and method that includes a first output voice adjustment unit to set a maximum volume level for each sound source and a voice synthesis unit to synthesize these adjusted voices, ensuring balanced output levels across all sound sources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple sound sources are synthesized and distributed without volume adjustment, then the content distribution is simple and fast, but the volume balance between sound sources becomes poor and some voices cannot be heard

Engineering Contradiction:
Improvevolume balanceVSAvoidprocessing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies preliminary action by adjusting the volume of each sound source before synthesis occurs. The adjustment unit sets appropriate volume levels for each sound source in advance, ensuring that when the sounds are synthesized and distributed, the volume balance is already optimized. This prevents the problem of some voices being inaudible due to poor volume balance while maintaining efficient processing.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If volume adjustment processing is added to each sound source, then the volume balance is improved, but the processing time and complexity increase

Engineering Contradiction:
Improvevolume balanceVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The volume adjustment is performed as a preliminary step before synthesis, allowing for optimized volume balance without adding significant processing time to the overall system. By setting volumes in advance rather than adjusting them during or after synthesis, the process becomes more efficient.

Inventive Principle:
Principle #10Preliminary action

3Ease of operation

If the voice of one sound source is made louder, then that sound source is more prominent, but other sound sources become difficult to hear

Engineering Contradiction:
Improveclarity of individual sound sourceVSAvoidaudibility of all sound sources
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent applies parameter changes by adjusting the volume parameter of each sound source individually. The adjustment unit modifies the volume levels of multiple sound sources simultaneously to achieve an optimal balance, ensuring that each sound source remains audible while maintaining appropriate prominence. This resolves the contradiction by finding the right parameter settings for all sound sources.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4383751B1Information processing device, information processing method, and program
Publication Date: 2026.02.18 SONY GROUP CORP
  • EP4383751B1 patent drawingFigure 1
  • EP4383751B1 patent drawingFigure 2
  • EP4383751B1 patent drawingFigure 3

AI summary

A device and a method are provided that adjust voices of a plurality of sound sources included in distribution content from an information processing device and make it easier to hear sound of each sound source by a reception terminal that receives and reproduces distribution content. A first output voice adjustment unit executes adjustment processing on an output voice of each of the plurality of sound sources, and content including synthesized voice data obtained by synthesizing the output voices corresponding to the sound sources adjusted by the first output voice adjustment unit is output. The first output voice adjustment unit executes output voice adjustment processing for matching a maximum value of a volume level corresponding to a frequency of the output voice of each sound source to a target level. Moreover, a second output voice adjustment unit executes the output voice adjustment processing in accordance with a type or a scene of the content.