Secure Text-to-Speech Synthesis with Voiceprint Authentication

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing text-to-speech synthesis systems in portable electronic devices lack security features to ensure that messages are read in the natural voice of the sender, and there is a need for secure conversion of text-based content to audio format while preventing unauthorized use of the sender's voice.

Innovation Solution

A method and system for secure text-to-speech synthesis that authenticates the originator or recipient of text-based content using digital signatures and voiceprints, allowing conversion to audio format only in accordance with the sender's voiceprint, ensuring secure and natural voice synthesis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If text-to-speech synthesis is implemented in portable electronic devices, then the ability to convert text to speech is improved, but security vulnerabilities arise that allow unauthorized use of the sender's voice

Engineering Contradiction:
Improvetext-to-speech conversion capabilityVSAvoidvoice authentication security
Core Design Contradiction:
Adaptability or versatilityVSReliability

Solution Approach 1:

The system performs preliminary authentication of the originator using digital signatures before allowing text-to-speech conversion. The voiceprint is verified against the originator's identity in advance, ensuring that only authenticated users can convert their signed text to speech, thereby preventing unauthorized voice synthesis

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary authentication mechanism that mediates between the text content and the voice synthesis process. The digital signature and voiceprint verification act as intermediaries to ensure that the text being converted to speech is genuinely from the claimed originator, adding a security layer without preventing the TTS functionality

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If voiceprint authentication is implemented, then security against unauthorized voice synthesis is improved, but system complexity increases

Engineering Contradiction:
Improvevoiceprint authentication securityVSAvoidauthentication system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent leverages existing universal authentication infrastructure (digital signatures and public key cryptography) that is already widely implemented in portable electronic devices. By building upon these existing security mechanisms rather than creating entirely new authentication protocols, the system achieves voiceprint authentication without proportionally increasing device complexity

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system uses digital signatures as a copyable and transferable authentication mechanism that can be verified without requiring the original private key. This allows secure authentication to be implemented through verifiable copies of the originator's signature, simplifying the authentication process while maintaining security

Inventive Principle:
Principle #26Copying

Data Source

PatentUS9166977B2Secure text-to-speech synthesis in portable electronic devices
Publication Date: 2015.10.20 MALIKIE INNOVATIONS LTD
  • US9166977B2 patent drawing
  • US9166977B2 patent drawing
  • US9166977B2 patent drawing

AI summary

A method for secure text-to-speech conversion of text using speech or voice synthesis that prevents the originator's voice from being used or distributed inappropriately or in an unauthorized manner is described. Security controls authenticate the sender of the message, and optionally the recipient, and ensure that the message is read in the originator's voice, not the voice of another person. Such controls permit an originator's voiceprint file to be publicly accessible, but limit its use for voice synthesis to text-based content created by the sender, or sent to a trusted recipient. In this way a person can be assured that their voice cannot be used for content they did not write.