Text-to-Speech Conversion for Voice Chat Accessibility

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing multiplayer game systems cannot accommodate players with speaking impairments or those in quiet environments, limiting their ability to participate in voice chat during multiplayer sessions.

Innovation Solution

Implementing a text-to-speech (TTS) conversion feature that allows players to input text, which is then converted to synthesized voice data, enabling them to communicate with others in the session without needing to speak aloud, with options for voice pitch selection and customizable text entry interfaces.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If voice chat is used for multiplayer communication, then communication capability is improved, but accessibility for players with speaking impairments deteriorates

Engineering Contradiction:
Improvecommunication capabilityVSAvoidaccessibility for players with speaking impairments
Core Design Contradiction:
Ease of operationVSAdaptability or versatility

Solution Approach 1:

The patent introduces a text-to-speech conversion service as an intermediary between the player and the voice chat system. Players input text through a text entry interface, and the conversion service translates it into synthesized speech that is transmitted through the existing voice chat infrastructure. This mediator enables players with speaking impairments to use the voice chat functionality without requiring direct vocal output.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If text-to-speech conversion is implemented, then accessibility for players with speaking impairments is improved, but device complexity increases

Engineering Contradiction:
Improveaccessibility for players with speaking impairmentsVSAvoidsystem complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent extracts the text-to-speech conversion functionality from the core voice chat system and implements it as a separate, independent service. This conversion service operates externally and interfaces with the voice chat system through standardized protocols, allowing the complexity of TTS processing to be isolated from the main communication infrastructure. Other players and systems interact with the voice chat without needing to aware of or accommodate the TTS conversion mechanism.

Inventive Principle:
Principle #2Taking out (Extraction)

3Adaptability or versatility

If synthesized voice is used instead of real voice, then privacy protection for players is improved, but audio quality and naturalness deteriorate

Engineering Contradiction:
Improveprivacy protectionVSAvoidaudio quality
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent utilizes advanced text-to-speech technology that dynamically adjusts multiple audio parameters to enhance the quality of synthesized speech. The system can modify pitch, tone, speed, and other acoustic characteristics to make the synthesized voice sound more natural and less distinguishable from real human speech. This parameter optimization maintains privacy protection while minimizing the audible difference between synthesized and real voices.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10839787B2Session text-to-speech conversion
Publication Date: 2020.11.17 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10839787B2 patent drawing
  • US10839787B2 patent drawing
  • US10839787B2 patent drawing

AI summary

Examples described herein provide various devices that enable users to participate in a multiplayer session. The examples allow a user that is unable to speak, or that is incapable of speaking, to participate in an in-session voice chat by inputting text and having the text converted to speech (e.g., synthesized voice data) that can then be sent to other devices participating in the session. The user enables a text-to-speech conversion feature on his or her own device. Based on the enabled feature, functionality enabling text to be entered is activated and the entered text is converted into speech data.