Automatic Spoken Word Emphasis via Predictive Pitch and Duration

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Non-professional speakers recording spoken content face challenges in maintaining high-quality speech characteristics, such as emphasis, tone, and diction, leading to monotony and difficulty in engaging audiences, with limited post-recording adjustment options.

Innovation Solution

The development of predictive models for pitch, duration, and spectral balance that analyze contextual and lexical information to automatically emphasize words in spoken content, allowing for post-recording emphasis adjustments without requiring users to re-record.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If a non-professional speaker records spoken content themselves, then production cost and complexity are reduced, but speech quality characteristics (emphasis, tone, diction) deteriorate leading to monotony

Engineering Contradiction:
Improveproduction simplicityVSAvoidspeech quality
Core Design Contradiction:
Ease of manufactureVSManufacturing precision

Solution Approach 1:

The patent introduces an automated emphasis system as an intermediary between the non-professional speaker and the final audio output. The system analyzes the recorded speech, identifies words that should be emphasized based on contextual and lexical information, and applies pitch and duration modifications automatically. This mediator compensates for the speaker's lack of professional voice acting skills while preserving the simplicity of self-recording.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system enables non-professional speakers to self-correct their own recordings without requiring professional voice actors or manual re-recording. The automated emphasis detection and application process allows the speaker to improve their own speech quality characteristics independently, making the system self-sufficient and eliminating the need for external professional intervention.

Inventive Principle:
Principle #25Self-service

2Ease of operation

If visual cues are provided to guide user performance, then ease of operation is improved, but reliability deteriorates because users may perform incorrectly according to the cues

Engineering Contradiction:
Improveguidance availabilityVSAvoidperformance accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

Instead of guiding the user on how to perform emphasis correctly during recording, the system inverts the approach by analyzing the recorded performance and automatically applying corrections. Rather than telling the user what to do, the system does it for them by detecting emphasis opportunities from contextual and lexical information and applying pitch/duration modifications automatically.

Inventive Principle:
Principle #13The other way round (Inversion)

Solution Approach 2:

The system allows the recording to serve itself by automatically detecting and correcting emphasis issues without requiring the user to interpret and follow visual cues. The automated analysis of contextual and lexical information enables the system to self-correct the performance, eliminating the reliability issues associated with user interpretation of guidance cues.

Inventive Principle:
Principle #25Self-service

3Manufacturing precision

If professional voice actors are used, then speech quality characteristics are improved, but production cost and complexity increase

Engineering Contradiction:
Improvespeech qualityVSAvoidproduction complexity
Core Design Contradiction:
Manufacturing precisionVSEase of manufacture

Solution Approach 1:

The patent replaces the mechanical system of professional voice acting with an automated computational system. Instead of relying on human voice actors to manually apply emphasis, pitch variation, and diction control, the system uses automated detection algorithms that analyze contextual and lexical information to determine emphasis requirements, then applies modifications programmatically. This substitution eliminates the need for professional voice actors while maintaining speech quality characteristics.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

4Manufacturing precision

If re-recording is performed to correct emphasis, then speech quality is improved, but loss of time increases

Engineering Contradiction:
Improveemphasis qualityVSAvoidrecording time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The system performs preliminary analysis of the recorded speech to identify words that should be emphasized based on contextual and lexical information before applying any modifications. By detecting emphasis requirements in advance and applying pitch/duration modifications automatically to the original recording, the system eliminates the need for time-consuming re-recording sessions while maintaining emphasis quality.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of discarding the original recording and requiring re-recording, the system recovers and enhances the original recording by applying automated emphasis modifications. The original performance is preserved and improved upon through selective pitch and duration adjustments, eliminating the time loss associated with discarding and re-recording while maintaining or improving emphasis quality.

Inventive Principle:
Principle #34Discarding and recovering

Data Source

PatentUS9852743B2Automatic emphasis of spoken words
Publication Date: 2017.12.26 ADOBE INC
  • US9852743B2 patent drawing
  • US9852743B2 patent drawing
  • US9852743B2 patent drawing

AI summary

Embodiments of the present invention provide systems, methods, and computer storage media directed towards automatic emphasis of spoken words. In one embodiment, a process may begin by identifying, within an audio recording, a word that is to be emphasized. Once identified, contextual and lexical information relating to the emphasized word can be extracted from the audio recording. This contextual and lexical information can be utilized in conjunction with a predictive model to determine a set of emphasis parameters for the identified word. These emphasis parameters can then be applied to the identified word to cause the word to be emphasized. Other embodiments may be described and/or claimed.