Oscillation Signal Generation for Voice Inflection in Video Playback

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies face challenges in effectively conveying voice inflection and vocalization speed to deaf and hard-of-hearing individuals during video playback, as text-based captions struggle to represent these aspects accurately.

Innovation Solution

An information processing apparatus that generates oscillation signals based on sound data waveforms to represent sound effects and vocalizations, using caption information and sound information from video files, allowing for separate and distinct representation of sound effects and vocalizations through oscillation, which are then provided to users via oscillation devices.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If text-based captions are used to represent sound content, then deaf and hard-of-hearing people can access video content, but the inflection, volume, and vocalization speed of voice cannot be represented

Engineering Contradiction:
Improvevoice inflection and vocalization speed informationVSAvoidcaption system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent segments caption data into two distinct types: sound-effect caption data (for non-vocal sounds like explosions) and vocalization caption data (for human speech). This segmentation allows each type to be processed differently, with vocalization captions being converted into oscillation signals that preserve voice inflection and speed information, while sound-effect captions remain as text descriptions.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces oscillation signals as an intermediary between the caption system and the user. These oscillation signals are generated from the waveform of sound data corresponding to vocalization captions, serving as a mediator that conveys voice characteristics (inflection, volume, speed) through tactile feedback, thereby preserving information that would otherwise be lost in text-based captions.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If automated haptification algorithm is used to generate haptic effects for sound effects, then haptic feedback is provided, but voice inflection and vocalization speed cannot be conveyed

Engineering Contradiction:
Improvehaptic feedback capabilityVSAvoidvocalization characteristics
Core Design Contradiction:
Ease of operationVSLoss of information

Solution Approach 1:

The patent applies different processing methods to different types of caption data based on their local quality requirements. Sound-effect captions receive automated haptification processing for tactile feedback, while vocalization captions receive oscillation signal generation that specifically preserves voice characteristics. This localized quality approach ensures each caption type gets the appropriate treatment for its specific information needs.

Inventive Principle:
Principle #3Local quality

3Device complexity

If caption data is not divided into sound-effect and vocalization types, then processing is simpler, but oscillation signals cannot be generated for vocalizations

Engineering Contradiction:
Improvecaption processing complexityVSAvoidvocalization information
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent performs preliminary classification of caption data into sound-effect and vocalization types before generating oscillation signals. This preliminary action (categorization) is performed on the caption information before the oscillation generation process, allowing the system to identify which captions should be converted into oscillation signals while keeping the overall processing framework relatively simple.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12101576B2Information processing apparatus, information processing method, and program
Publication Date: 2024.09.24 SONY GROUP CORP
  • US12101576B2 patent drawing
  • US12101576B2 patent drawing
  • US12101576B2 patent drawing

AI summary

There is provided an information processing apparatus, an information processing method, and a program that make it possible to assist deaf and hard-of-hearing people in viewing a video when the video is being played back. The information processing apparatus includes a controller. The controller generates at least one of an oscillation signal corresponding to sound-effect caption data or an oscillation signal corresponding to vocalization caption data on the basis of a waveform of sound data using a result of analyzing caption information and sound information that are included in a video file, the sound-effect caption data being used to represent a sound effect in the form of text information, the vocalization caption data being used to represent a vocalization of a person in the form of text information, the sound-effect caption data and the vocalization caption data being included in caption data that is included in the caption information, the sound data being included in the sound information.