Multidimensional Audio Positioning for Text-to-Speech Scanning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-speech technologies require users to listen to each item of textual content individually, making it inefficient for users to find interesting topics among multiple aggregated content sources, especially with the increasing number of online content providers and the need for improved multidimensional audio output.

Innovation Solution

A method and system that generate a simultaneous stream of multiple audio selections from text passages converted to speech, with distinct speech attributes to position each selection in a multidimensional sound space, allowing users to efficiently scan and focus on specific content using posture movements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If text to speech conversion is applied to multiple content sources simultaneously, then user efficiency in finding interesting topics is improved, but audio source differentiation and selection become difficult

Engineering Contradiction:
Improveuser efficiency in finding interesting topicsVSAvoidaudio source differentiation and selection
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent applies spatial dimensionality to audio output by positioning multiple speech sources in a three-dimensional sound space. Each content source is assigned a specific spatial location, allowing users to scan through content by moving their head or adjusting listening direction, analogous to visual scanning but in an auditory spatial environment.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent applies different speech attributes (pitch, rate, timbre) to different audio sources based on their content characteristics and spatial positions. This local differentiation of quality attributes enables users to distinguish between multiple simultaneous speech sources and selectively focus on interesting content.

Inventive Principle:
Principle #3Local quality

2Productivity

If multiple audio selections are played simultaneously in a multidimensional sound space, then content scanning efficiency is improved, but posture tracking and selection precision become challenging

Engineering Contradiction:
Improvecontent scanning efficiencyVSAvoidposture tracking and selection precision
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent incorporates posture tracking feedback mechanisms that monitor user head position and movement, and use this information to dynamically adjust audio source positioning and volume. This feedback loop enables precise selection of desired content based on user posture while maintaining efficient simultaneous playback of multiple sources.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS11356792B2Selecting a primary source of text to speech based on posture
Publication Date: 2022.06.07 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11356792B2 patent drawing
  • US11356792B2 patent drawing
  • US11356792B2 patent drawing

AI summary

A computer converts each content sources from textual content to speech comprising a separate audio selection. The computer applies, to each audio selection, one or more speech attributes to specify the audio attributes that select a respective position of the respective audio selection from among multiple positions within a multidimensional sound space and audibly distinguish one or more characteristics of the respective audio selection from other audio selections, wherein the respective position of the respective audio selection reflects a rank of the respective audio selection as ordered by interest to a user. The computer outputs a simultaneous stream of the multiple audio selections to an audio output device for stereo play of the audio selections within the multiple positions within the multidimensional sound space to the user, with the multiple positions reflecting the content sources ordered by interest.