CGR Speech Synthesis via Language Characteristic Modification

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In virtual speech synthesis, existing technologies struggle to adapt speech inputs to match the language characteristics and relationships of participants in a computer-generated reality (CGR) environment, leading to potential misuse of inappropriate speech patterns.

Innovation Solution

A device that includes a display, an audio sensor, and processors to modify speech inputs based on language characteristic values associated with a fictional character, ensuring that the synthesized CGR speech aligns with the intended audience and context within the CGR environment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If speech inputs are directly transmitted without modification in CGR environments, then communication efficiency is improved, but speech appropriateness and realism deteriorate

Engineering Contradiction:
Improvecommunication efficiencyVSAvoidspeech appropriateness
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system performs preliminary analysis of the fictional character's language characteristics (accent, tone, vocabulary, grammar patterns) before the speech transmission occurs. This pre-processing enables the modification of user speech to match the character's linguistic profile in advance, ensuring appropriateness without compromising communication efficiency.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary speech modification system that acts as a mediator between the user's original speech and the transmitted speech. This intermediary layer analyzes the fictional character's language characteristics and transforms the user's speech accordingly, maintaining both communication efficiency and speech appropriateness.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If speech modification based on language characteristics is implemented, then speech realism and appropriateness are improved, but system complexity increases

Engineering Contradiction:
Improvespeech realismVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The speech modification system is segmented into distinct functional modules: a language characteristic analysis module that extracts accent, tone, and vocabulary patterns from fictional character data, and a speech transformation module that applies these characteristics to user speech. This segmentation reduces overall system complexity by making each module independently manageable.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system creates a linguistic model (copy) of the fictional character's speech patterns by analyzing their language characteristics. This copied linguistic profile is then applied to modify user speech, avoiding the need for complex real-time generation of entirely new speech patterns while maintaining realism.

Inventive Principle:
Principle #26Copying

3Manufacturing precision

If comprehensive language characteristic analysis is performed, then speech accuracy to character profile is improved, but processing time increases

Engineering Contradiction:
Improvespeech accuracyVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs comprehensive language characteristic analysis in advance, before speech transmission is needed. By pre-processing and storing the linguistic profile of the fictional character (accent, tone, vocabulary preferences, grammar patterns), the system avoids time-consuming analysis during actual speech communication, thus maintaining high accuracy while reducing processing time.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The speech modification system dynamically adjusts the level of analysis based on context. For routine communications, it applies pre-analyzed language characteristics quickly. For more critical or nuanced interactions, it performs more comprehensive real-time analysis, optimizing the balance between accuracy and processing time.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12223943B2Assisted speech
Publication Date: 2025.02.11 APPLE INC
  • US12223943B2 patent drawing
  • US12223943B2 patent drawing
  • US12223943B2 patent drawing

AI summary

Various implementations disclosed herein include devices, systems, and methods for synthesizing virtual speech. In various implementations, a device includes a display, an audio sensor, a non-transitory memory and one or more processors coupled with the non-transitory memory. A computer-generated reality (CGR) representation of a fictional character is displayed in a CGR environment on the display. A speech input is received from a first person via the audio sensor. The speech input is modified based on one or more language characteristic values associated with the fictional character in order to generate CGR speech. The CGR speech is outputted in the CGR environment via the CGR representation of the fictional character.