Voice Classification via Resonant Frequency Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice classification technologies face challenges in accurately distinguishing between individuals based on voice samples, particularly in family settings where permission settings on devices need to be adapted differently for various users.
Innovation Solution
A method and controller that generate a unique user profile by analyzing the resonant frequencies in voice samples, using a machine learning algorithm to classify users into specific classes such as father, mother, or child, by extracting features like frequency, amplitude, and position in the frequency spectrum, and adapting permission settings accordingly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If voice classification is performed using traditional methods, then the system is simple to implement, but the classification accuracy is insufficient to distinguish between individuals in family settings
Solution Approach 1:
The voice classification system segments the voice signal into multiple frequency bands and extracts formant frequencies (F1, F2, F3, F4) as distinct features. This segmentation of the spectral content enables precise identification of individual voice characteristics, resolving the contradiction between simple implementation and high classification accuracy by breaking down the complex voice signal into manageable spectral components.
Solution Approach 2:
The system transitions from analyzing single-dimensional voice features to multi-dimensional spectral analysis by examining formant frequencies across different frequency bands. This dimensional expansion in the frequency domain provides richer discriminatory information for distinguishing between family members, achieving high classification accuracy without excessive system complexity.
2Adaptability or versatility
If generic permission settings are used for all users, then the device is easy to operate, but it cannot provide personalized access control for different family members
Solution Approach 1:
The voice classification system enables self-service personalization by automatically identifying users through their voice characteristics and applying appropriate permission settings without manual configuration. The system extracts formant frequencies and compares them against stored profiles to autonomously determine user identity and grant corresponding access rights, achieving personalization while maintaining operational simplicity.
Solution Approach 2:
The system performs preliminary voice profiling during a setup phase where family members record their voices for classification. This preliminary action creates stored voice profiles that enable automatic, personalized permission assignment during subsequent use, eliminating the need for manual configuration and providing both adaptability and ease of operation.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Enables high-accuracy user classification and adaptive permission settings on devices, ensuring secure and personalized access based on user profiles.
Implementation Method 1
The body's cavities cause the sound waves produced when a person speaks to vibrate with higher amplitudes at certain frequencies (i.e. resonant frequencies). The resonant frequencies are dependent on the shape and size of a cavity and therefore, in general, a person's voice will have resonant frequencies unique to that person.
Data Source
AI summary
A controller and method of classifying a user into one of a plurality of user classes. One or more voice samples are received from the user, from which a frequency spectrum is generated. One or more values defining respective features of the frequency spectrum are extracted from the frequency spectrum. Each of the respective features are defined by values of frequency, amplitude, and/or position in the spectrum. One or more of the respective features are resonant frequencies in the voice of the user. A user profile of the user is generated and comprises the extracted one or more values. The user profile is supplied to a machine learning algorithm that is trained to classify users as belonging to one of the plurality of user classes based on the one or more values in their respective user profile.


