Automatic Spoken Word Emphasis via Predictive Pitch and Duration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Non-professional speakers recording spoken content face challenges in maintaining high-quality speech characteristics, such as emphasis, tone, and diction, leading to monotony and difficulty in engaging audiences, with limited post-recording adjustment options.
Innovation Solution
The development of predictive models for pitch, duration, and spectral balance that analyze contextual and lexical information to automatically emphasize words in spoken content, allowing for post-recording emphasis adjustments without requiring users to re-record.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If a non-professional speaker records spoken content themselves, then production cost and complexity are reduced, but speech quality characteristics (emphasis, tone, diction) deteriorate leading to monotony
Solution Approach 1:
The patent introduces an automated emphasis system as an intermediary between the non-professional speaker and the final audio output. The system analyzes the recorded speech, identifies words that should be emphasized based on contextual and lexical information, and applies pitch and duration modifications automatically. This mediator compensates for the speaker's lack of professional voice acting skills while preserving the simplicity of self-recording.
Solution Approach 2:
The system enables non-professional speakers to self-correct their own recordings without requiring professional voice actors or manual re-recording. The automated emphasis detection and application process allows the speaker to improve their own speech quality characteristics independently, making the system self-sufficient and eliminating the need for external professional intervention.
2Ease of operation
If visual cues are provided to guide user performance, then ease of operation is improved, but reliability deteriorates because users may perform incorrectly according to the cues
Solution Approach 1:
Instead of guiding the user on how to perform emphasis correctly during recording, the system inverts the approach by analyzing the recorded performance and automatically applying corrections. Rather than telling the user what to do, the system does it for them by detecting emphasis opportunities from contextual and lexical information and applying pitch/duration modifications automatically.
Solution Approach 2:
The system allows the recording to serve itself by automatically detecting and correcting emphasis issues without requiring the user to interpret and follow visual cues. The automated analysis of contextual and lexical information enables the system to self-correct the performance, eliminating the reliability issues associated with user interpretation of guidance cues.
3Manufacturing precision
If professional voice actors are used, then speech quality characteristics are improved, but production cost and complexity increase
Solution Approach 1:
The patent replaces the mechanical system of professional voice acting with an automated computational system. Instead of relying on human voice actors to manually apply emphasis, pitch variation, and diction control, the system uses automated detection algorithms that analyze contextual and lexical information to determine emphasis requirements, then applies modifications programmatically. This substitution eliminates the need for professional voice actors while maintaining speech quality characteristics.
4Manufacturing precision
If re-recording is performed to correct emphasis, then speech quality is improved, but loss of time increases
Solution Approach 1:
The system performs preliminary analysis of the recorded speech to identify words that should be emphasized based on contextual and lexical information before applying any modifications. By detecting emphasis requirements in advance and applying pitch/duration modifications automatically to the original recording, the system eliminates the need for time-consuming re-recording sessions while maintaining emphasis quality.
Solution Approach 2:
Instead of discarding the original recording and requiring re-recording, the system recovers and enhances the original recording by applying automated emphasis modifications. The original performance is preserved and improved upon through selective pitch and duration adjustments, eliminating the time loss associated with discarding and re-recording while maintaining or improving emphasis quality.
Data Source
AI summary
Embodiments of the present invention provide systems, methods, and computer storage media directed towards automatic emphasis of spoken words. In one embodiment, a process may begin by identifying, within an audio recording, a word that is to be emphasized. Once identified, contextual and lexical information relating to the emphasized word can be extracted from the audio recording. This contextual and lexical information can be utilized in conjunction with a predictive model to determine a set of emphasis parameters for the identified word. These emphasis parameters can then be applied to the identified word to cause the word to be emphasized. Other embodiments may be described and/or claimed.


