Voice Bot Synthetic Responses for Voice Privacy Protection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies lack regulations to prevent misuse of voice recordings and fail to protect user privacy in scenarios where voice data is collected without consent, particularly in unsolicited calls and spam calls that can record a person's voice without permission.

Innovation Solution

A device is configured to allow users to select and render text-based labels for synthesized voice responses, using synthetic speech generators to respond to calls without exposing their actual voice, with features like auto-muting the microphone for unknown callers and providing ranked or user-specified message lists.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If users respond to calls using their own voice, then communication effectiveness is improved, but voice privacy and security are compromised due to unauthorized recording

Engineering Contradiction:
Improvecommunication effectivenessVSAvoidvoice privacy exposure
Core Design Contradiction:
Ease of operationVSObject-affected harmful factors

Solution Approach 1:

The patent creates a synthetic copy of the user's voice that mimics their vocal characteristics without using their actual voice recordings. The system generates artificial voice responses that sound like the user but are completely synthesized, allowing communication while preventing exposure of the user's real voice data to callers or third parties.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent introduces a synthetic voice generation system as an intermediary between the user and the caller. Instead of the user's actual voice being transmitted, the system acts as a mediator that converts text inputs into synthetic speech that resembles the user's voice, thereby protecting the user's voice privacy while maintaining communication flow.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Object-affected harmful factors

If synthesized voice responses are used to protect privacy, then voice recording security is improved, but communication naturalness may deteriorate

Engineering Contradiction:
Improvevoice recording securityVSAvoidcommunication naturalness
Core Design Contradiction:
Object-affected harmful factorsVSEase of operation

Solution Approach 1:

The patent adjusts various parameters of the synthetic voice generation including pitch, tone, speech rate, and vocal timbre to match the user's natural speaking characteristics. By fine-tuning these acoustic parameters, the system creates synthetic speech that closely resembles the user's natural voice, making the communication feel more authentic and natural despite being synthesized.

Inventive Principle:
Principle #35Parameter changes

3Loss of energy

If text-based synthesized responses are implemented, then bandwidth usage is reduced, but response customization flexibility is limited

Engineering Contradiction:
Improvebandwidth usageVSAvoidresponse customization
Core Design Contradiction:
Loss of energyVSAdaptability or versatility

Solution Approach 1:

The patent implements a universal text-to-speech system that can generate multiple types of responses (polite refusals, busy messages, do-not-disturb notifications, etc.) from a single text-based interface. The system handles various communication scenarios using the same synthetic voice generation mechanism, providing diverse response options while maintaining efficient bandwidth usage through text-based input rather than audio recording.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12363224B2Voice bot for service providers
Publication Date: 2025.07.15 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12363224B2 patent drawing
  • US12363224B2 patent drawing
  • US12363224B2 patent drawing

AI summary

A device is configured to communicate on a mobile communications network. An incoming call is received, and it is determined that the incoming call meets a predetermined criteria indicating a probable source of the incoming call. On a display of the device, an option is rendered for answering the incoming call with a generated voice response in lieu of a voice of a user of the device. Text options for generating a voice response are also rendered. The incoming call is answered and generated speech corresponding to the selected text option is sent.