Graphical Vocalization Adjustment System
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional computing systems are unable to effectively assist users in learning to alter their pronunciation to match predetermined standards, as they struggle to provide clear and actionable feedback on physical adjustments needed to improve vocalization qualities.
Innovation Solution
A method and system that gather audio and spatial data of a user's vocalization, analyze the data to identify incorrect qualities, and provide graphical representations of suggested facial adjustments in real-time to help the user change their vocalization to match preferred qualities, using a controller that generates vector diagrams and augmented reality to guide the user in making necessary mouth and tongue positions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If conventional computing systems provide feedback on pronunciation, then users can receive guidance on improving vocalization, but the feedback is not clear or actionable enough to guide physical adjustments
Solution Approach 1:
The system uses color-coded visual indicators to represent different pronunciation qualities and required adjustments. Facial feature targets are displayed with color cues indicating the direction and magnitude of adjustment needed, making abstract audio feedback concrete and immediately understandable.
Solution Approach 2:
The system creates a virtual overlay copy of the user's face showing ideal facial positions superimposed on the actual face. This visual copy allows users to directly compare their current pronunciation posture with the target posture, providing intuitive guidance without requiring complex technical understanding.
2Manufacturing precision
If the system provides detailed feedback on multiple facial elements, then users can achieve precise pronunciation improvement, but the complexity of the system increases
Solution Approach 1:
The system breaks down pronunciation into discrete facial elements (lips, teeth, tongue, jaw) and provides independent visual targets for each. This segmentation allows precise control over multiple aspects of pronunciation while presenting information in manageable, separate units rather than as an overwhelming whole.
Solution Approach 2:
The system introduces an intermediary visual layer between the audio input and the user's physical adjustments. This intermediary graphical interface translates complex audio analysis into simple visual targets that guide facial movements, reducing the perceived complexity while maintaining precision.
3Loss of information
If the system provides comprehensive visual feedback, then users can learn pronunciation effectively, but the amount of information presented overwhelms the user
Solution Approach 1:
The system applies different visual qualities and levels of detail to different facial elements based on their importance and the user's current performance. Elements requiring adjustment are highlighted with greater visual prominence, while elements already in correct position receive minimal or no visual feedback, allowing users to focus on critical areas without being overwhelmed by comprehensive but equally weighted information.
Data Source
AI summary
Audio of a user speaking is gathered while spatial data of the face of the user is gathered. Positions of elements of the face of are identified, wherein relative positions of the elements cause a plurality of qualities of the user voice. A subset of positions of the elements are identified to have caused a detected first quality of the user voice during the period of time. Alternate positions of the one or more elements are identified that are determined to cause the user voice to have a second quality rather than the first quality. A graphical representation of the face that depicts one or more adjustments from the subset of the positions to the alternate positions is provided to the user.


