Dynamic Audio Commentary Generation via Sound Element Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio production technologies limit audio file creation to pre-recorded sound contents, resulting in monotonous game commentaries in platforms like game entertainment systems, failing to incorporate real-time player voices and emotions effectively.

Innovation Solution

An information processing apparatus and method that selects sound elements related to scene features, establishes a correspondence relationship between these elements and features, and generates customized personalized sound based on a correspondence relationship library, allowing real-time voice incorporation and dynamic commentary creation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If pre-recorded commentary audio files are used in game platforms, then audio production is simplified and standardized, but the commentary becomes monotonous and fails to engage users

Engineering Contradiction:
Improveaudio production simplicityVSAvoidcommentary personalization
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent segments a complete commentary audio into multiple independent sound elements (e.g., player introduction, action description, scoring announcement). Each sound element can be independently selected and recombined based on different game scenarios and player preferences, enabling personalized commentary while maintaining production efficiency through modular design

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent merges multiple pre-recorded sound elements into dynamic commentary compositions. By combining selected sound elements with real-time game data and player-specific attributes, the system generates customized commentary that adapts to different users while leveraging the efficiency of pre-produced audio assets

Inventive Principle:
Principle #5Merging (Combining)

2Loss of time

If pre-recorded sound contents are used for audio file creation, then production time and complexity are reduced, but the audio content lacks real-time player voice and emotion integration

Engineering Contradiction:
Improveaudio production timeVSAvoidplayer voice and emotion information
Core Design Contradiction:
Loss of timeVSLoss of information

Solution Approach 1:

The patent introduces sound elements as intermediary components that bridge pre-recorded audio and real-time player information. These sound elements serve as building blocks that can be populated with player-specific data (voices, emotions, statistics) while maintaining the structural efficiency of pre-produced audio content

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transforms static pre-recorded audio into dynamic commentary by enabling real-time selection and customization of sound elements based on player attributes, game state, and user preferences. This allows the system to adapt audio content dynamically without requiring full real-time recording and processing

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS11417315B2Information processing apparatus and information processing method and computer-readable storage medium
Publication Date: 2022.08.16 SONY GROUP CORP
  • US11417315B2 patent drawing
  • US11417315B2 patent drawing
  • US11417315B2 patent drawing

AI summary

An information processing apparatus and an information processing method as well as a computer readable storage medium are provided. The information processing apparatus includes a processing circuitry configured to: select, from a sound, sound elements which are related to scene features during making of the sound; establish a correspondence relationship including a first correspondence relationship between the scene features and the sound elements and between the respective sound elements, and store the scene features and the sound elements as well as the correspondence relationship in association in a correspondence relationship library; and generate, based on a reproduction scene feature and the correspondence relationship library, a sound to be reproduced.