Systems and methods for voice evaluation

The AI-based voice evaluation system addresses the limitations of current tools by allowing self-assessment and remote monitoring, enhancing accessibility and precision in diagnosing and treating voice disorders.

WO2026080579A9PCT designated stage Publication Date: 2026-05-21SONOVOICE INC +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SONOVOICE INC
Filing Date
2025-10-08
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Current voice evaluation tools require expensive equipment and trained professionals, are inaccessible to the general public, and rely on inaccurate self-reporting, leading to delayed or incorrect diagnoses and treatments for voice-related disorders.

Method used

An artificial intelligence-based voice evaluation system that is hardware agnostic, allowing users to self-assess and monitor their vocal health, providing real-time analysis and personalized recommendations for treatment and rehabilitation.

Benefits of technology

Improves accessibility and precision in voice disorder diagnosis and treatment by enabling self-evaluation and remote monitoring, offering real-time feedback and personalized therapeutic approaches.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure US2025050020_21052026_PF_FP_ABST
    Figure US2025050020_21052026_PF_FP_ABST
Patent Text Reader

Abstract

An artificial intelligence based voice evaluation system designed to improve the diagnosis and treatment of vocal disorders. The artificial intelligence based voice evaluation system utilizes machine learning to provide real-time analysis of aerodynamic and acoustic data corresponding to user vocal health. The artificial intelligence based voice evaluation system is operable to monitor a vocal health of a user and to generate at least one recommendation based on the user vocal health data.
Need to check novelty before this filing date? Find Prior Art

Description

Attorney Docket No.: 1532 / 3 PCTSYSTEMS AND METHODS FOR VOICE EVALUATIONCROSS-REFERENCE TO RELATED APPLICATIONS

[0001] This application claims priority to and the benefit of U.S. Provisional Patent Application No. 63 / 704,834, filed on October 08, 2024, the entire content of which is herein incorporated by reference in its entirety.BACKGROUND

[0002] One in eight adults experience voice-related issues (e.g., dysphonia) each year in the United States, at an estimated cost of nearly fifteen billion dollars and considerable negative effects on quality of life and mental well-being. If detected and treated early, voice-related issues can be properly treated, thereby improving the quality of life of an affected individual. However, there are significant barriers to obtaining treatment for voice evaluation and treatment. Many clinicians that can properly evaluate and treat voice-related problems are located in academic medical centers which are not easily accessible to the general public. Further, the diagnosis and treatment can be expensive and unaffordable for individuals that suffer from voice-related issues (e.g., teachers, singers, performers).FIELD OF THE DISCLOSURE

[0003] The present disclosure generally relates to systems and methods for voice evaluation and more specifically to hardware agnostic systems and methods for assessing the vocal health and function of users using artificial intelligence.DESCRIPTION OF RELATED ART

[0004] Current voice evaluation tools require expensive equipment connected to a desktop computer and a speech-language pathologist (SLP) trained in perceptual evaluation and interpretation of vocal function data. Additionally, current voice evaluation tools are reliant on the individual seeking voice rehabilitation to self-report their vocal health via questionnaires and to monitor their vocal health during rehabilitation regimens without tools that generate longitudinal datasets, particularly with the broader shift to telehealth. Self-reporting can result in inaccurate information that may result in an incorrect diagnosis or a delay in treatment. Therefore, there is a need for a hardware agnostic voice evaluation system that improves the access to healthcare for the millions of people that suffer from vocal-related disorders to improve diagnosis and treatment.Attorney Docket No.: 1532 / 3 PCTBRIEF SUMMARY

[0005] The present disclosure includes an artificial intelligence based voice evaluation system designed to aid the diagnosis and treatment of voice disorders by (1) improving the precision with which an individual’s vocal function ability and dysfunction can be measured and (2) giving affected individuals the ability to self-monitor and self-evaluate their vocal health without a physician or speech-language pathologist. The artificial intelligence based voice evaluation system improves detection of subtle voice abnormalities (e.g., decreased vibratory regularity, inefficient airflow). The artificial intelligence based voice evaluation system further provides personalized therapeutic approaches and recommendations to aid diagnosis and treatment.

[0006] The artificial intelligence based voice evaluation system provides a hardware agnostic system usable for healthcare professionals and patients. The artificial intelligence based voice evaluation system enables patients to assess and monitor their vocal health. Advantageously, the artificial intelligence based voice evaluation system enables remote access and analysis of vocal health and function data while providing a comprehensive assessment and diagnosis. The artificial intelligence based voice evaluation system includes an artificial intelligence component designed to automate voice evaluation. The artificial intelligence component includes machine learning (ML) based voice training guidance and feedback for tracking and optimizing the vocal health and function of a user. Yet another advantage of the artificial intelligence based voice evaluation system is real-time or near real-time analysis of continuous voice health data.

[0007] In some embodiments, the artificial intelligence based voice evaluation system is designed to determine a physiological condition of a user. The physiological condition includes determining at least one vocal status and / or condition. The artificial intelligence based voice evaluation system is further operable to determine a physiological condition of a user based on the vocal health of a user. The physiological condition includes a voice-related condition and / or a non-voice related condition.

[0008] In some embodiments, the artificial intelligence based voice evaluation system is operable to generate a personalized recommendation for a user based on vocal health and function data. The personalized recommendation includes a treatment plan, a rehabilitation plan, a physician and / or speech-language pathologist, and other recommendations for addressing a condition of a user. The personalized recommendation can include a videoAttorney Docket No.: 1532 / 3 PCTconsultation with the recommended clinician and a voice exercise regimen. In some embodiments, the artificial intelligence based voice evaluation system is operable to track a user’s vocal health and condition during a rehabilitation program. Advantageously, the artificial intelligence based voice evaluation system is operable to update a user’s rehabilitation program based on the user’s progress. In some embodiments, the artificial intelligence based voice evaluation system is in network communication with at least one smart wearable device and / or an internet-of-things (loT) device.

[0009] In some embodiments, an artificial intelligence based voice evaluation system is disclosed. The artificial intelligence based voice evaluation system provides real-time or near real-time analysis of voice health data and voice function data for at least one user. The artificial intelligence based voice evaluation system is operable to monitor and analyze vocal health data and vocal function data corresponding to a plurality of users and to generate personalized recommendations based on the vocal health data.

[0010] The artificial intelligence based voice evaluation system is operable to receive vocal health and function data from at least one remote device and / or at least one third-party source to provide real-time analysis and recommendations. The remote device includes, but is not limited to, a computer, a laptop, a cell phone, a smart wearable device, a vocal analysis device, and other similar remote devices operable for capturing and transmitting vocal health and function data. Advantageously, the artificial intelligence based voice evaluation system is operable to optimize the personalized recommendations to improve the success of a user’s rehabilitation.

[0011] In some embodiments, an artificial intelligence based voice evaluation system including at least one remote device including a user interface, at least one remote server including a software platform including a plurality of dashboards, a calibration engine, an analytics engine, and at least one database is disclosed. The at least one remote device, the at least one remote server, the calibration engine, the analytics engine and the at least one database are in network communication. The at least one remote device is operable to capture vocal health and function data corresponding to at least one user. The user vocal health and function data is transmitted to the at least one remote server. After receiving the user vocal health and function data, the analytics engine is operable to analyze the user vocal health and function data to determine at least one vocal physiological condition of the user. Based on the vocal physiological condition of the user, the at least one remote server is operable to transmit the determined vocal physiological condition of the user to a remote device. The at least one remoteAttorney Docket No.: 1532 / 3 PCTdevice is further operable to display the determined vocal physiological condition of the user.

[0012] The at least one analytics engine is operable to determine at least one non- voice related physiological condition and at least one personalized recommendation based on the at least one determined vocal physiological condition of the user. The at least one personalized recommendation includes a treatment recommendation, a rehabilitation recommendation, and / or a physician recommendation.

[0013] In some embodiments, the artificial intelligence based vocal evaluation system includes at least one vocal evaluation device. The at least one vocal evaluation device is operable to capture sensor data (e.g., aerodynamic, acoustic) corresponding to a user performing vocal exercises. The artificial intelligence based vocal evaluation system is operable to calibrate the at least one vocal evaluation device based on a user. The at least one vocal evaluation device is operable for real-time capture and transmission of sensor data to a vocal evaluation software platform.

[0014] In at least one aspect of the present disclosure, an artificial intelligence based voice evaluation system for assisting diagnosis and treatment of voice disorders is disclosed. The artificial intelligence based voice evaluation system may include at least one vocal evaluation device including at least one sensor, at least one user interface, and an analytics engine. The at least one sensor may be configured to capture vocal data including vocal health data and vocal function data. The at least one vocal evaluation device can display the captured vocal data via the at least one user interface. The at least one vocal evaluation device can store the captured vocal data. The analytics engine can determine a vocal health of a user based on the captured vocal data.

[0015] The artificial intelligence based voice evaluation system may include a multi-modal dataset including aerodynamic data and acoustic data. The analytics engine may be further configured for generating at least one recommendation based on the determined user vocal health. The at least one recommendation may include a treatment recommendation, a rehabilitation recommendation, or a physician recommendation. The at least one sensor may include at least one of an audio sensor, an electroglottograph sensor, a humidity sensor, an imaging sensor, a movement sensor, an aerodynamic sensor, and / or a temperature sensor. The artificial intelligence based vocal evaluation system may further include a calibration engine that can generate at least one sensor recommendation based on the vocal data. The at least one sensor recommendation may include positioning, digital gain control, and / or signal filtering.Attorney Docket No.: 1532 / 3 PCTThe vocal data may include vocal imaging data, vocal aerodynamic data, vocal acoustic data, and / or self-evaluation data.BRIEF DESCRIPTION OF THE SEVERAL VIEWS OF THE DRAWINGS

[0016] The embodiments illustrated, described, and discussed herein are illustrative of the present disclosure. As these embodiments of the present disclosure are described with reference to illustrations, various modifications or adaptations of the methods and or specific structures described may become apparent to those skilled in the art. It will be appreciated that modifications and variations are covered by the above teachings and within the scope of the appended claims without departing from the spirit and intended scope thereof. All such modifications, adaptations, or variations that rely upon the teachings of the present disclosure, and through which these teachings have advanced the art, are considered to be within the spirit and scope of the present disclosure. Hence, these descriptions and drawings should not be considered in a limiting sense, as it is understood that the present invention is in no way limited to only the embodiments illustrated.

[0017] FIG. 1 illustrates a schematic diagram of an artificial intelligence based vocal evaluation system according to at least one aspect of the present disclosure.

[0018] FIG. 2 illustrates a schematic diagram of an artificial intelligence based vocal evaluation system according to at least one aspect of the present disclosure.

[0019] FIG. 3 illustrates a schematic diagram of an artificial intelligence based vocal evaluation system according to at least one aspect of the present disclosure.

[0020] FIG. 4 illustrates a flow diagram of machine learning model training for an artificial intelligence based vocal evaluation system according to at least one aspect of the present disclosure.

[0021] FIG. 5 illustrates a flow diagram for operating a vocal evaluation device and vocal evaluation platform of an artificial intelligence based vocal evaluation system at least one aspect of the present disclosure.

[0022] FIG. 6 illustrates a flow diagram for light emitting diode (LED) functionality of a vocal evaluation system according to at least one aspect of the present disclosure.

[0023] FIG. 7 illustrates a flow diagram for operation of an artificial intelligence based vocal evaluation system according to at least one aspect of the present disclosure.

[0024] FIG. 8 illustrates a flow diagram for setup and calibration of an artificial intelligenceAttorney Docket No.: 1532 / 3 PCTbased vocal evaluation system according to at least one aspect of the present disclosure.

[0025] FIG. 9 illustrates a flow diagram for operation of an artificial intelligence vocal evaluation system according to one at least one aspect of the present disclosure.

[0026] FIG. 10 illustrates a flow diagram for operation of a vocal evaluation system according to at least one aspect of the present disclosure.

[0027] FIG. 11 illustrates a schematic diagram of an artificial intelligence based vocal evaluation system according to at least one aspect of the present disclosure.

[0028] FIG. 12 illustrates a schematic diagram of a remote server of an artificial intelligence based vocal evaluation system according to at least one aspect of the present disclosure.

[0029] FIG. 13 illustrates a schematic diagram of a computer of an artificial intelligence based vocal evaluation system according to at least one aspect of the present disclosure.

[0030] FIG. 14 illustrates a schematic diagram of an artificial intelligence based vocal evaluation system according to at least one aspect of the present disclosure.

[0031] FIG. 15 illustrates a schematic diagram of an artificial intelligence based vocal evaluation system according to at least one aspect of the present disclosure.DETAILED DESCRIPTION

[0032] For the purposes of promoting an understanding of the present disclosure, reference will be made to preferred embodiments and specific language will be used to describe the same. It will nevertheless be understood that no limitation of the scope of the disclosure is thereby intended, such alteration and further modifications of the disclosure as illustrated herein, being contemplated as would normally occur to one skilled in the art to which the disclosure relates.

[0033] Articles “a” and “an” are used herein to refer to one or to more than one (i.e., at least one) of the grammatical object of the article. By way of example, “a composite” means at least one composite and can include more than one composite.

[0034] Throughout the specification, the terms “about” and / or “approximately” may be used in conjunction with numerical values and / or ranges. The term “about” is understood to mean those values near to a recited value. For example, “about 40 [units]”may mean within + / - 25% of 40 (e.g., from 30 to 50), within + / - 20%, + / - 15%, + / - 10%, + / - 9%, + / -8 %, + / - 7%, + / - 6%, + / - 5%, + / - 4%, + / - 3%, + / -2 %, + / - 1%, less than + / - 1%, or any other value or range ofAttorney Docket No.: 1532 / 3 PCTvalues therein or there below. Furthermore, the phrases “less than about [a value]” or “greater than about [a value]” should be understood in view of the definition of the term "about" provided herein. The terms "about" and "approximately" may be used interchangeably.

[0035] As used herein, the verb “comprise” as is used in this description and in the claims and its conjugations are used in its non-limiting sense to mean that items following the word are included, but items not specifically mentioned are not excluded.

[0036] Throughout the specification the word “comprising,” or variations such as “comprises” or “comprising,” will be understood to imply the inclusion of a stated element, integer or step, or group of elements, integers, or steps, but not the exclusion of any other element, integer or step, or group of elements, integers, or steps. The present disclosure may suitably “comprise”, “consist of’, or “consist essentially of’, the steps, elements, and / or reagents described in the claims.

[0037] It is further noted that the claims may be drafted to exclude any optional element. As such, this statement is intended to serve as antecedent basis for use of such exclusive terminology as “solely”, “only”, and the like in connection with the recitation of claim elements, or the use of a “negative” limitation.

[0038] Unless defined otherwise, all technical and scientific terms used herein have the same meanings as commonly understood by one of ordinary skill in the art to which this disclosure belongs. Preferred methods, devices, and materials are described, although any methods and materials similar or equivalent to those described herein can be used in the practice or testing of the present disclosure. All references cited herein are incorporated by reference in their entirety.

[0039] The subject matter described herein includes an artificial intelligence based voice evaluation system. The artificial intelligence based voice evaluation system is hardware agnostic and improves the portability and precision of voice evaluation. The artificial intelligence based voice evaluation system utilizes a synchronous multi-modal dataset of aerodynamic and acoustic parameters. The aerodynamic and acoustic parameters can be captured during a user’s performance of semi-occluded vocal tract (SOVT) exercises across a range of aerodynamic resistances. The artificial intelligence based voice evaluation system is operable to correlate user vocal data with laryngeal pathology and other related physiological conditions. For example, and without limitation, the user vocal data includes vocal health data and vocal function data. Advantageously, the artificial intelligence based voice evaluationAttorney Docket No.: 1532 / 3 PCTsystem is operable to provide a real-time evaluation of user vocal health and function. In some embodiments, the artificial intelligence based voice evaluation system includes a machine learning component designed for automatic analysis of the multi-modal dataset of aerodynamic and acoustic parameters. The artificial intelligence based voice evaluation system is designed for at least one user and a plurality of users. The artificial intelligence based voice evaluation system provides personalized vocal recommendations including treatment recommendations, rehabilitation recommendations, and voice-specialized clinician recommendations (including laryngologists and speech-language pathologists).

[0040] In some aspects, an artificial intelligence based voice evaluation system including at least one remote device including a user interface and at least one remote server including a software platform including a plurality of dashboards, an analytics engine, and at least one database is disclosed. The at least one remote device and the at least one remote server are in network communication. The at least one remote device is operable to capture vocal data (e.g., aerodynamic data, acoustic data). In some aspects, the vocal data includes at least one biomarker. Alternatively, or additionally, the vocal data is captured by the software platform via a third-party source (e.g., web crawler). After receiving the vocal data, the analytics engine of the at least one remote server is designed to analyze the vocal data to determine a real-time status of the vocal health and function of a user. Based on the user vocal health and function, the at least one remote server is further operable to provide at least one vocal health recommendation and / or at least one vocal function recommendation.

[0041] In some embodiments, the artificial intelligence based voice evaluation system is in network communication with at least one wearable device. The at least one wearable device includes, but is not limited to, a smart wearable device, an Intemet-of-Things (loT) device, and other devices operable to capture a physiological condition of a user.

[0042] In some aspects, the at least one database includes user vocal health data, user vocal function data, vocal health standards, and medical data for a plurality of users. In some aspects, the analytics engine is operable to determine at least one trend based on the vocal health data. Based on the at least one trend, the analytics engine is operable to identify one or more vocal health conditions, disorders, and / or disorders related to the user vocal health data. The analytics engine is further operable to generate a vocal recommendation. In some embodiments, the vocal recommendation includes an amount of sleep, a medication recommendation, a nutrition recommendation, a vocal exercise, and other recommendations for improving a condition of a user.Attorney Docket No.: 1532 / 3 PCT

[0043] In some embodiments, the at least one software platform includes a messaging platform. The messaging platform is operable for electronic communication between each remote device of the plurality of remote devices. In some embodiments, the at least one vocal health recommendation further includes at least one voice-specialized clinician recommendation (laryngologist and / or speech-language pathologist). The software platform is further operable to automatically create an electronic messaging conversation between at least one remote device corresponding to the at least one voice-specialized clinician recommendation (laryngologist and / or speech-language pathologist) clinician recommendation and the at least one user.

[0044] FIG. 1 is a block diagram of an artificial intelligence based vocal evaluation system 100 according to one embodiment of the present invention. The artificial intelligence based vocal evaluation system 100 includes a plurality of vocal sensors 102, a remote device 104 with local storage 106, and a remote server 108 with a database 110. Advantageously, the artificial intelligence based vocal evaluation system is designed to capture and / or receive vocal data of at least one user and to provide real-time feedback, analysis, and recommendations based on the user vocal data. For example, and without limitation, the user vocal data includes vocal health data and vocal function data. In some embodiments, the artificial intelligence based vocal evaluation system is operable to determine a user’s vocal output, phonatory function, and respiratory physiological conditions.

[0045] The plurality of vocal sensors 102 includes an audio sensor 112, an electroglottograph sensor 114, a humidity sensor 115, an imaging sensor 116, a movement sensor 118, an aerodynamic sensor 120, and a temperature sensor 121.

[0046] The audio sensor 112 is operable to detect and capture audio data (e.g., speech and voice) corresponding to at least one user. For example, and without limitation, the audio sensor 112 is operable to detect and measure sound waves and / or acoustic signals. The audio sensor 112 includes but is not limited to an omnidirectional microphone, a unidirectional microphone, a condenser microphone, a frequency sensor, an acoustic sensor, a piezoelectric sensor, and / or other sensors operable to capture audio data. In some embodiments, the audio sensor includes a preamplifier and an analog-to-digital converter for improving the capture of audio data.

[0047] The electroglottograph sensor 114 is operable for noninvasive measurement of a degree of contact of vibrating vocal folds when performing speech. The electroglottograph sensor is operable to measure variations in electrical impedance across the vocal folds as they vibrate. The artificial intelligence based vocal evaluation system is operable to correlate anAttorney Docket No.: 1532 / 3 PCTamplitude of the EGG signal with an amount of vocal fold closure.

[0048] The humidity sensor 115 is operable to measure the relative humidity in an environment (e.g., air). For example, and without limitation, the humidity sensor includes a capacitive sensor, a resistive sensor, and / or a thermal sensor. The artificial intelligence based vocal evaluation system is operable to correlate environment humidity with user vocal health data.

[0049] In some embodiments, the imaging sensor 116 is operable to capture image data of the upper respiratory system (e.g., larynx) of a user. The imaging sensor is operable for videostroboscopy, stroboscopic visual-perceptual assessment, vibratory function analysis, and nonvibratory function analysis. In some embodiments, the imaging sensor is attached or built- in to the at least one remote device 104.

[0050] The movement sensor 118 is operable to detect and measure movement during speech exercises and / or tasks. In some embodiments, the movement sensor 118 includes an accelerometer, a gyroscope, a voice dosimeter, an inertial measurement unit, and other movement related sensors. In some embodiments, the movement sensor includes a movement sensor operable to detect and measure vibration of the vocal folds when a user is vocalizing.

[0051] In some embodiments, the aerodynamic sensor 120 is operable to determine respiratory conditions of a user. The respiratory conditions of a user includes a respiratory rate. In some embodiments, the aerodynamic sensor is operable to detect and measure airflow mass, airflow pressure, and other airflow measurements during voicing. For example, and without limitation, the aerodynamic sensor includes a pneumotachograph device including a differential pressure transducer to determine differential pressure, intraoral air pressure, and other similar respiratory measurements. For further example, the pneumotachograph device includes a Fleisch pneumotachograph or Lilly pneumotachograph.

[0052] In some embodiments, the aerodynamic sensor 120 includes a pulse oximeter for monitoring oxygen saturation. In some embodiments, the pulse oximeter is a photoplethysmogram sensor. The aerodynamic sensor includes a respiratory effort transducer for measuring the change in thoracic or abdominal circumference. In some embodiments, the aerodynamic sensor includes a pressure sensor and / or an airflow sensor. For example, and without limitation, the pressure sensor is operable to capture the distribution of pressure in the glottis when a user is vocalizing. For example, and without limitation, the airflow sensor is operable to noninvasively capture a user’s airflow while vocalizing. The airflow sensor isAttorney Docket No.: 1532 / 3 PCToperable to determine the airflow over time while a user is vocalizing.

[0053] In some embodiments, the temperature sensor 121 is operable to determine a temperature of an environment and / or of a user while performing vocal exercise. For example, and without limitation, the temperature sensor includes a thermocouple, a resistance temperature detector, a negative temperature coefficient (NTC) thermistor, a semi-conductor based temperature sensor, or other similar temperature capturing devices. The temperature sensor measures and monitors changes in temperature. The artificial intelligence based vocal evaluation system is operable to correlate temperature data with the user vocal health and function data.

[0054] The remote device 104 can be a smartphone, a tablet, a laptop, a desktop computer, and other similar remote devices. The remote device 104 includes a processor 122, an analytics engine 124, and a control interface 126. The remote device 104 accepts data input from plurality of vocal sensors 102. The remote device 104 also accepts data input from the remote server 108. The remote device 104 stores data in a local storage 106.

[0055] The local storage 106 on the remote device 104 includes a user profile 128, historical user data 130, and vocal assessment data 132. The user profile 128 stores physiological information about the user, including but not limited to, age, weight, height, body mass and fat, hydration, gender, medical history (e.g., vocal diseases), fitness (e.g., fitness level), dietary restrictions and preferences, stress level, experience, and / or occupational information (e.g., occupation, shift information). The user profile further includes user-reported measures. For example, and without limitation, the user-reported measures include vocal health and quality of life outcome measures (e.g., Voice Handicap Index). The historical user data 130 includes information gathered from the plurality of vocal health sensors 102 and / or the at least one remote device 104.

[0056] The remote server 108 includes global historical acoustic data 134, global historical aerodynamic data 136, global historical user data 138, global recommendation data 140, global treatment data 142, and global vocal assessment data 144. The remote server 108 further includes a global analytics engine 146 and a calibration engine 148. The global historical acoustic data 134, the global historical aerodynamic data 136, and the global historical user data 138 include data from a plurality of users and / or vocal health devices and system.

[0057] The remote server 108 is operable to store vocal data and user profile data. The vocal data includes data (e.g., vocal health and function) captured by the plurality of vocalAttorney Docket No.: 1532 / 3 PCTsensors. The plurality of vocal sensors and the plurality of remote devices are operable to transmit the captured data to the remote server in real-time.

[0058] The global analytics engine 146 is designed to receive real-time physiological (e.g., vocal health and function) data and the user profile data. The global analytics engine 146 is configured to evaluate the condition of a user based on the real-time vocal health data, vocal function data, and the user profile data. The global analytics engine 146 is operable for acoustic analysis 150, aerodynamic analysis 152, imaging analysis 154, recommendation analysis 156, and treatment analysis 158.

[0059] Acoustic analysis 150 and aerodynamic analysis 152 includes analysis of user vocal health data to determine the vocal health of a user. The artificial intelligence based vocal evaluation system is operable to identify signs of potential diseases, disorders, or other conditions based on the user vocal health and function. The artificial intelligence based vocal evaluation system is further operable to identify the absence of biomarkers or physiological conditions corresponding to one or more diseases, disorders, or other conditions. In some embodiments, the global analytics engine is operable to determine a glottal airflow rate, an average interpolated air pressure, a mean vocal SPL (dB), and vocal frequency (Hz) during a vocal exercise. The global analytics engine is further operable to determine (1) intraoral pressure, (2) aerodynamic resistance, (3) rate of change of aerodynamic resistance, (4) SOVT exercises, (5) phonatory stress on vocal folds, (6) somatosensory vibrotactile feedback produced by the palate and the anterior oral cavity / mask, and (7) acoustic parameters (e.g., intensity, fundamental frequency, regularity, harmonicity) across aerodynamic resistances. The imaging analysis 154 includes analysis of image data corresponding to a user’s body to determine the vocal health and function of a user. The global analytics engine is operable to provide a recommendation for improving a user’s posture (e.g., vocal tract posture) and phonation based on the data received from the plurality of vocal health sensors. The recommendation analysis 156 includes analysis of a user’s vocal health to generate at least one recommendation. Advantageously, the global analytics engine is operable to determine vocal health and function similarities between users and to compare and generate new recommendations based on the change in vocal health and function for each user. The treatment analysis 158 includes the generation of a treatment plan based on a user’s vocal health and function and the evaluation of a user treatment plan.

[0060] Yet another advantage of the global analytics engine 146 is determining whether a user is at risk of an injury and generating an alert or command based on the risk. For example,Attorney Docket No.: 1532 / 3 PCTand without limitation, the global analytics engine can compare the real-time vocal data to the medical history of a user. The global analytics engine is designed to detect that a user’s heart rate is above a predetermined threshold for a vocal exercise. The threshold is based on the vocal health history of the user. Advantageously, the remote server is further operable to generate an alert for the user to contact a medical professional, to change a treatment plan, and / or modify a vocal exercise based on the real-time physiological data and vocal health data.

[0061] In some embodiments, the artificial intelligence based voice evaluation system is configured to monitor a user’s vocal health and function during an activity and to identify vocal health and function changes. Advantageously, the artificial intelligence based vocal evaluation system generates at least one alert and at least one recommendation based on the user’s vocal health and function and changes in the user’s vocal health and function. For example, and without limitation, the artificial intelligence based vocal evaluation system is operable to determine whether the phonation threshold pressure is straining or stressing a user’s vocal folds. If the artificial intelligence based vocal evaluation system determines that the phonation threshold pressure is too high, the artificial intelligence based vocal evaluation system is operable to generate an alert indicating to avoid generating the sound corresponding to the high phonation threshold pressure. Advantageously, the artificial intelligence based vocal evaluation system is operable to adjust alerts based on real-time user vocal health data.

[0062] The artificial intelligence based voice evaluation system is configured to receive vocal health and function standards data. Advantageously, the artificial intelligence based voice evaluation system is designed to receive real-time vocal health and function standards updates from government agencies and vocal health organizations and to update vocal health diagnosis, analysis, and recommendations based on the updated vocal health and function standards.

[0063] In some embodiments, the artificial based voice evaluation system is operable for voice evaluation analysis involving physiological data, biomechanical data, and aerodynamic data. Vocal evaluation analysis includes auditory-perceptual analysis, imaging (e.g., laryngeal endoscopic) analysis, acoustic analysis, aerodynamic analysis, and self-evaluation. Stroboscopic visual-perceptual assessment is used to determine regularity, vocal fold vibratory amplitude, mucosal waves, vocal fold phase symmetry, vertical level, and glottal closure patterns. Laryngeal visualization analysis further includes vibratory and non-vibratory function. The nonvibratory function includes vocal fold abduction and adduction.

[0064] In some embodiments, the artificial intelligence based vocal evaluation system is operable to receive imaging data corresponding to user speech. For example, and withoutAttorney Docket No.: 1532 / 3 PCTlimitation, the imaging data includes images of vocal fold structure and function. The artificial intelligence based vocal evaluation system is operable to assess the vocal fold medial edges, vocal fold mobility, supraglottic activity during phonation, laryngeal maneuvers, vocal fold vibratory amplitude, mucosal wave, vocal fold phase symmetry, vertical level, glottal closure pattern, vocal fold abduction, and vocal fold abduction.

[0065] The vocal fold edges analysis includes a rating of the edges of a vocal fold. For example, and without limitation, the vocal folds edges analysis includes analyzing the vocal folds in an abducted position. The vocal fold edges rating can include smooth, straight, bowed, irregular, or rough. Vocal fold mobility identifies the movement of each vocal fold toward and away from a midline when producing laryngeal adduction / abduction and maximum adduction / abduction tasks. A vocal fold mobility rating can include normal, reduced, or absent. Supraglottic activity includes a degree of compression of supraglottic structures during sustained phonation. Supraglottic compression can include unilateral medial compression, bilateral medial compression, anteroposterior compression, and sphincteric compression. The type of compression can be expressed as mild, moderate, or severe. In some embodiments, the vocal fold edges analysis is operable to generate a vocal fold health and function score.

[0066] In some embodiments, the artificial intelligence based vocal evaluation system is operable to track the frequency and phase of vocal fold vibration. Amplitude is a lateral movement of a vibration portion of a vocal cord in the medial plane during phonation. Rating of amplitude includes evaluation of medial to lateral excursion of the midmembranous position of a vocal fold during phonation ranging from about zero percent to about one hundred percent. One hundred percent corresponds to a total width of vocal folds. Mucosal wave is a lateral movement of the mucosa over the vocal fold body. The mucosal wave can be rated based on the mucosal wave movement from the medial edge toward the lateral surface of the vocal fold in increments of 25%, ranging from 0% to 100%. 100% is the total visible width of the vocal fold.

[0067] In some embodiments, the artificial intelligence based vocal evaluation system is operable to determine and monitor left / right phase symmetry. The left / right phase symmetry includes a degree to which the vocal folds appear as mirror images of each other during an apparent (videostroboscopic) glottal cycle in the timing of opening, closing, and maximum lateral-medial excursion.

[0068] In some embodiments, the artificial intelligence based vocal evaluation system is operable to determine and monitor vertical level. Vertical level is a difference in a vertical plane between two vocal folds during a maximum closed phase of a glottal cycle. In someAttorney Docket No.: 1532 / 3 PCTembodiments, the artificial intelligence based vocal evaluation system is operable to determine glottal closure pattern. Glottal closure pattern includes complete closure (e.g., no apparent gap on a maximal closure), anterior gaps, irregular closures, degrees of closure along a length of the vocal folds, spindle-shaped gap, posterior gap, hourglass gaps, an absence of closure, and variable closure.

[0069] In some embodiments, the artificial intelligence based vocal evaluation system is operable to receive auditory-perceptual data. The auditory-perceptual data includes but is not limited to overall severity, breathiness, roughness, strain, pitch, loudness, weakness, resonance characteristics, tremor, spasms, and other vocal quality data.

[0070] In some embodiments, the artificial intelligence based voice evaluation system is operable to measure noise in voice signal data. For example, and without limitation, the noise measurement is based on levels of periodic energy and / or aperiodic energy in the voice acoustic data during sustained vowels, speech, and other vocal exercises. For further example, in some embodiments, the artificial intelligence based voice evaluation system is operable to analyze an entire range of dysphonia using cepstral based measurements. In some embodiments, the measurements include jitter, shimmer, and other variations. In some embodiments, a vocal CPP is used to measure an amplitude of a peak in a cepstral based measurement for sustained vowels, speech, and related voice exercises.

[0071] In some embodiments, the artificial intelligence based voice evaluation system is designed for aerodynamic assessment of vocal health data. The aerodynamic assessment can include glottal aerodynamic measures. The glottal aerodynamic measures include but are not limited to glottal airflow rate and subglottal air pressure. In some embodiments, the artificial intelligence based voice evaluation system further determines intraoral air pressure. In some embodiments, the aerodynamic assessment includes vocal loudness (e.g., sound pressure level), pitch (fo), quality, habitual vocal SPL, minimum vocal SPL, mean focal fo, minimum fo, maximum fo, and vocal cepstral peak prominence.

[0072] In some embodiments, the artificial intelligence based vocal evaluation system is operable to suggest one or more vocal-related tasks and / or exercises for determining the vocal health and / or function of a user. For example, and without limitation, the vocal-related tasks and / or exercises includes (1) breathing cycles, (2) repeated inhalation at a desired rate, (3) phonation at a desired pitch and / or loudness, (4) phonation at variable pitches, (5) variable phonation at a constant loudness, (6) sustain vowel duration, (7) reading a passage, (8) loudness range, (9) pitch range, (10) SOVTE, and other vocal-related tasks and / or exercise for determining a user’ s vocal health.Attorney Docket No.: 1532 / 3 PCT

[0073] The calibration engine 148 is operable for calibrating voice recording data and recording systems to measure SPL values. The calibration engine 148 is further operable to alter the settings of at least one vocal sensor to improve the capture of the user data. For example, and without limitation, the calibration engine 148 is operable to determine that captured audio data cannot be analyzed because the audio is too quiet. In response, the calibration engine is operable to recommend adjusting a position of an audio device (e.g., microphone) to improve capture of audio data. For further example, and without limitation, the calibration engine is operable for digital gain control and signal filtering (e.g., high pass filter, low pass filter) to improve audio capture.

[0074] The plurality of vocal sensors 102, the remote device 104 with local storage 106, and the remote server 108 can connect directly (e.g., Universal Serial Bus (USB) or equivalent) or wirelessly (e.g., Bluetooth®, Wi-Fi®, ZigBee®) through systems designed to exchange data between various data collection sources. In some embodiments, the plurality of vocal sensors 102, the remote device 104, and the remote server 108 are in network communication.

[0075] In some embodiments, the artificial intelligence based vocal evaluation system includes an artificial intelligence engine 160 operable to receive real-time or near real-time data from the plurality of vocal sensors 102 and the at least one remote device 104. The artificial intelligence engine can generate a user vocal score based on the received vocal sensor data. The at least one remote device is operable to receive user input for physiological user information. The artificial intelligence engine is operable to generate a user's real-time or near real-time user vocal health and function score. Additionally, the artificial intelligence engine is operable to correlate the real-time or near real-time user vocal data with the global historical data, the global historical aerodynamic data, the global historical user data, the global recommendation data, the global treatment data, and the global vocal assessment data. The artificial intelligence engine is operable to use historical data to provide prompts and / or questions to a user to improve analysis. Additionally, the artificial intelligence engine is operable to provide user-specific questions based on the user profile, the historical user data, and user vocal assessment data.

[0076] In at least one aspect, the artificial intelligence engine is operable to generate a realtime or near real-time alert based on the real-time or near real-time physiological data, user vocal health data, user vocal function data, the global historical data, the global historical aerodynamic data, the global historical user data, the global recommendation data, the global treatment data, and the global vocal assessment data. The at least one alert includes changes inAttorney Docket No.: 1532 / 3 PCTvocal health data (e.g., instability in vocal fold vibration), vocal function data, and other similar alerts relating to the voice evaluation. Additionally, the artificial intelligence based vocal evaluation system is configured to provide a recommendation based on the alert. For example, and without limitation, the artificial intelligence engine is operable to determine that a user’s vocal folds are likely experiencing increased frequency and force of collision when a user is working. The artificial intelligence engine is operable to suggest a user refraining from their work or to limit the use of the user’s vocal folds. For further example, and without limitation, the artificial intelligence engine is operable to determine a lack of vocal quality improvement when using an SOVT posture. As a result of the lack of vocal quality improvement, the artificial intelligence engine is operable to determine that a user has vibratory impairments.

[0077] For example, and without limitation, in some aspects, the artificial intelligence engine includes at least one machine learning algorithm. The machine learning algorithm is operable to receive a multi-modal dataset including aerodynamic and acoustic parameters corresponding to vocal data. The multi-modal dataset can further include imaging data corresponding to a user. For example, and without limitation, the imaging data includes a plurality of images corresponding to a user while vocalizing. The machine learning algorithm is operable to quantify SOVT functionality based on the multi-modal dataset. For example, and without limitation, the machine learning algorithm is operable to estimate a level of vocal functionality based on the multi-modal dataset. The machine learning algorithm is operable to generate a baseline condition of a user and to modify the baseline condition based on user data including but not limited to age, gender, occupation, and other user data.

[0078] The machine learning algorithm is operable to receive and analyze acoustic data including shimmer, pitch sigma, the ratio of spectral energy, and the ratio of the actual amplitude of the CPP to the expected amplitude as determined via linear regression. The machine learning algorithm is operable to classify user into target groups based on user vocal health data. The machine learning algorithm is operable to receive aerodynamic variables including mean flow rate (MFR), mean vocal SPL and / captured from a sensor and / or remote device for classification. The machine learning algorithm includes multinomial logistic regression, naive bayes classifications, and random forests to classify users. In some embodiments, the machine learning algorithm is refined based on feature importance. For example, and without limitation, the feature selection for random forest can be performed using feature importance, permutation importance, and recursive feature elimination (RFE). For naive bayes classification, feature selection can be performed via mutual information and a chi-Attorney Docket No.: 1532 / 3 PCTsquared test. Feature selection for multinomial logistic regression can be performed using information from ANOVAs, recursive feature elimination, and Lasso regression. In some embodiments, the artificial intelligence engine includes hidden Markov models (HMMs), gaussian mixture models (GMMs), support vector machines (SVMs), and convolutional neural networks (CNNs).

[0079] In some embodiments, the multi-modal dataset includes vocal intensity, vocal fundamental frequency, and cepstral peak prominence. The user health data includes hydration, self-reported vocal fatigue, and experience. In some embodiments, the artificial intelligence based voice evaluation system is designed to suggest at least one physical condition of a user based on the user vocal data. For example, and without limitation, the at least one physical condition includes a cold, a flu, strep throat, a headache, and other medical conditions.

[0080] The artificial intelligence engine is operable to generate a vocal evaluation instruction for at least one user. For example, and without limitation, the at least one vocal evaluation instruction includes raising vocal pitch to a higher than typical level. The artificial intelligence engine includes a visualization component. The visualization component is operable to display the captured and analyzed real-time or near real-time data. For example, and without limitation, the artificial intelligence is operable to generate a virtual representation of a user’s vocal folds based on the user vocal data.

[0081] FIG. 2 illustrates a schematic of a machine learning vocal evaluation platform of an artificial intelligence based voice evaluation system according to one embodiment of the present invention. The vocal evaluation platform 202 includes a machine learning voice evaluation engine 204, a vocal health database 206, voice health metadata 208, a healthcare report component 210, and a physiological intake (e.g., questionnaire) component 212. The machine learning voice evaluation engine 204 receives vocal health data from the vocal health database 206, vocal health metadata 208, one or more reports via a remote device and / or third party corresponding to the reporting component 210, and user profile information from the physiological intake component 212. The vocal health metadata 208 includes information that improves the vocal health and wellbeing of an individual. The healthcare report component 210 includes medical reports provided by healthcare facilities and physicians.

[0082] FIG. 3 illustrates a flow diagram of an artificial intelligence based voice evaluation system. The artificial intelligence based voice evaluation system includes a remote device 302, at least one user identifying cloud 304, at least one data cloud 306, and at least one softwareAttorney Docket No.: 1532 / 3 PCTplatform 308. The remote device 302, the at least one user identifying cloud 304, the at least one data cloud 306, and the at least one software platform 308 are in network communication. For example, and without limitation, the artificial intelligence based voice evaluation system includes a federated database with at least two cloud platforms. At least one cloud platform is operable to identify data and at least one cloud platform is for additional data. Each cloud platform has at least one database and an application programming interface. In some embodiments, the at least two cloud platforms do not share any access points.

[0083] The at least one software platform 308 requires a user account before granting access to user data. The user account is created using at least two tokens. At least one token provides access to user identifying information and a second token provides access to other data. In some embodiments, the at least one software platform includes an administrative layer, a clinical layer, and a user layer. The administrative layer is operable to create, read, update, and delete clinician data, create and update patient data, and manage billing operations. The clinician layer is operable to read and update clinician data, and create, read, update, and delete patient data. The user layer is operable to create, read, update, and delete user data corresponding to a verified user. The user identifying cloud includes user identifying data (e.g., name, demographics, device ID, personal preferences). The data cloud includes anonymized user generated data and data generated by the software platform and / or user that does not identify a user.

[0084] As shown in FIG. 3, the at least one remote device 302 is operable for device pairing 310, device pairing confirmation 312, initiating a device test 314, and transmitting test results 316. The at least one user identifying cloud 304 is operable to receive a login API request 318 from the at least one software platform 308. The at least one user identifying cloud 304 is further operable to generate a login API response 320 (e.g., login successful). The at least one data cloud 306 is operable to receive a plurality of outputs from the at least one software platform 308.

[0085] The at least one software platform 308 includes a login page 322 and is operable to receive manual input (e.g., username and password) 324. The at least one software platform 308 is operable to receive a login API response from the user identifying cloud 304. The Login API response includes a successful login 326. Based on the successful login 326, the at least one software platform 308 is operable to generate a home page 328. The at least one software platform is operable to transmit a data API request 330 to the at least one data cloud 308 and to receive data required to run the software application 332.Attorney Docket No.: 1532 / 3 PCT

[0086] The at least one software platform 308 is further operable to receive user input including a survey request 344. In response, the at least one software platform generates a survey page 346. The at least one software platform 308 is operable to transmit a survey request 348 to the data cloud 306 and to receive a survey 350 from the cloud. The at least one software platform is operable to receive user input corresponding to the survey 352. The at least one software platform 308 is operable to display the survey input 354 and to transmit the survey results 356 to the data cloud 306.

[0087] The at least one software platform 308 is operable to receive user input requesting a device test 334 and to display a device testing page 336. The at least one software platform is operable to transmit a device pairing request 310 and to receive a device pairing confirmation 312 from the at least one remote device. After receiving the device pairing confirmation 312, the at least one software platform 308 receives user input for a device test 336. The at least one software platform is operable to initiate a test 314 and to receive the test results 316 from the at least one remote device. The at least one software platform is operable to display the test results 340 and to transmit the test results 340 to the cloud 306.

[0088] FIG. 4 illustrates a flow diagram of machine learning model training process of an artificial intelligence based vocal evaluation system. The machine learning model training process includes (1) data collection 402, (2) preprocessing the collected data 404, (3) model selection 406, (4) model training 408, (5) performance evaluation 410, and (6) model deployment 412. For example, and without limitation, the data collection includes receiving vocal data (e.g., vocal health and vocal function data) and annotating the vocal data (e.g., vocal health status). The collected vocal data is modified by simulating different environmental conditions and adding noise to create diverse samples. For further example, and without limitation, at least 100 data samples are created. The preprocessing step includes reducing noise, normalizing data, extracting features, and diving a dataset into training data, validation data, and tests sets for future evaluation. The model selection step includes applying a plurality of models and monitoring the model’s improvement. For example, and without limitation, a hidden Markov model is used when the collected data includes between 100 and 1000 data points. In some embodiments, a gaussian mixture model is used when the collected data includes at least 100 data points per mixture component. In some embodiments, the machine learning model includes a support vector machine when the collected data includes between 100 and 1000 data points per class. In some embodiments, the machine learning model includes at least one convolutional neural network with the collected data includes at least 10,000 dataAttorney Docket No.: 1532 / 3 PCTpoints. The machine learning model training includes turning hyperparameter and k-fold cross- validation. The machine learning model is operable to evaluate the vocal data based on desired metrics and / or threshold. After a machine learning model is selected, the machine learning model is integrated into the software platform.

[0089] In some embodiments, the artificial intelligence based vocal evaluation system includes at least one vocal evaluation device. For example, and without limitation, the at least one vocal evaluation device is operable to capture aerodynamic and acoustic data from oral and vocal exercises. The at least one vocal evaluation device includes at least one mechanical mechanism (e.g., tube, slider) and at least one sensor. In some embodiments, the at least one vocal evaluation device includes (1) at least two pressure sensors, (2) at least one acoustic microphone, (3) at least one humidity sensor, (4) at least one temperature sensor, and (5) at least one location sensor for a mechanical closure mechanism that restricts airflow. The at least one vocal evaluation device further includes at least one microprocessor. For example, and without limitation, the at least one microprocessor is operable to read the captured sensor data received at about 11 KHz. The at least one vocal evaluation device further includes a storage component (e.g., memory) for storing the captured aerodynamic and acoustic data. The at least one vocal evaluation device includes at least one indicator (e.g., light emitting diode (LED), at least one power supply component (e.g., battery, universal serial bus (USB) connection). The at least one vocal evaluation device is operable for wired and / or wireless communication to transmit the captured data in real-time. In some embodiments, the at least one vocal evaluation device is operable to analyze the captured data to determine at least one physiological condition of a user. For example, and without limitation, the at least one vocal evaluation device is operable to determine the vocal health and function of a user based on the captured aerodynamic and acoustic data.

[0090] The least one vocal evaluation device is operable for network communication with a vocal evaluation platform. In some embodiments, the at least one vocal evaluation device includes a slider component. The slider component is operable to move along a track and is manually adjustable. The at least one vocal evaluation device is operable for a setup mode. During setup mode, a first temperature, humidity and slider position is captured and transmitted to the vocal evaluation platform. The vocal evaluation platform is operable to receive user input and to adjust the gain on the microphone of the vocal evaluation device. The vocal evaluation device is operable to undergo adjustment exercises based on a user to configure the slider component, microphone gain, and other components of the vocal evaluation device. The atAttorney Docket No.: 1532 / 3 PCTleast one vocal evaluation device is operable to receive a selected vocal evaluation test from the vocal evaluation platform and to capture respiratory and acoustic data corresponding to the selected vocal evaluation test. The captured data includes elapsed time, audio signal, pressure signals, and other similar data. The captured data is stored in a storage component (e.g., memory) of the vocal evaluation device and is transferable to a remote device and the vocal evaluation platform.

[0091] FIG. 5 illustrates a flow diagram for operating a vocal evaluation device and vocal evaluation platform according to at least one aspect of the present disclosure. FIG. 6 illustrates a flow diagram for LED functionality according to at least one aspect of the present disclosure. FIG. 7 illustrates a flow diagram for operation of a vocal evaluation system according to at least one aspect of the present disclosure. FIG. 8 illustrates a flow diagram for setup and calibration of a vocal evaluation system according to at least one aspect of the present disclosure. FIG. 9 illustrates a flow diagram for operation of a vocal evaluation system according to at least one aspect of the present disclosure. FIG. 10 illustrates a flow diagram for operation of a vocal evaluation system according to at least one aspect of the present disclosure.

[0092] FIG. 11 depicts a system diagram 4500 illustrating a client / server architecture in accordance with embodiments of the present disclosure. The server application 4502 is configured to provide a video application and mobile application for an artificial intelligence based vocal evaluation system. A server application 4502 is hosted on a remote server 4504 within a cloud computing environment 4506. The server application 4502 is provided on a non- transitory computer-readable medium including a plurality of machine-readable instructions, which when executed by one or more processors of the server 4504, are adapted to cause the server 4504 to generate the video platform and mobile application.

[0093] The server application 4502 is configured to communicate over a network 4508. In a preferred embodiment, the network 4508 is the Internet. In other embodiments, the network 4508 may be restricted to a private local area network (LAN) and / or private wide area network (WAN). The network 4508 provides connectivity with a plurality of client devices including a personal computer 4510 hosting a client application 4512, and a mobile device 4514 hosting a mobile app 4516. The network 4508 also provides connectivity for an Internet-Of-Things (loT) device 4518 hosting an loT application 4520, and to back-end services 4522. Advantageously, the back-end services are operable to communicate with third-party application programming interfaces (APIs) to either provide or receive data that can be used by the system to provide recommendations. Third-party applications provide algorithms for analysis of data. The back-Attorney Docket No.: 1532 / 3 PCTend services may provide data gathered within the artificial intelligence based vocal evaluation system through the third-party APIs and receives results from the algorithms provided back to the back-end services to provide further recommendations or take further actions within the artificial intelligence based voice evaluation system.

[0094] FIG. 12 depicts a block diagram 4600 of the server 4504 of FIG. 11 for hosting at least a portion of the server application 4502 of FIG. 11 in accordance with embodiments of the present disclosure. The server 4504 may be any of the hardware servers referenced in this disclosure. The server 4504 may include at least one of a processor 4602, a main memory 4604, a database 4606, a datacenter network interface 4608, and an administration user interface (UI) 4610. The server 4504 may be configured to host one or more virtualized servers. For example, the virtual server may be an Ubuntu® server or the like. The server 4504 may also be configured to host a virtual container. For example, the virtual server may be the DOCKER® virtual server or the like. In some embodiments, the virtual server and or virtual container may be distributed over a plurality of hardware servers using hypervisor technology.

[0095] The processor 4602 may be a multi-core server class processor suitable for hardware virtualization. The processor 4602 may support at least a 64-bit architecture and a single instruction multiple data (SIMD) instruction set. The memory 4604 may include a combination of volatile memory (e.g., random access memory) and non-volatile memory (e.g., flash memory). The database 4606 may include one or more hard drives.

[0096] The datacenter network interface 4608 may provide one or more high-speed communication ports to the data center switches, routers, and / or network storage appliances. The datacenter network interface may include high-speed optical Ethernet, InfiniBand (IB), Internet Small Computer System Interface iSCSI, and / or Fibre Channel interfaces. The administration UI may support local and / or remote configuration of the server by a data center administrator.

[0097] FIG. 13 depicts a block diagram 4700 of the personal computer 4510 of FIG. 11 in accordance with embodiments of the present disclosure. The personal computer 4510 may be any of the devices referenced in this disclosure. The personal computer 4510 may include at least a processor 4702, a memory 4704, a display 4706, a user interface (UI) 4708, and a network interface 4710. The personal computer 4510 may include an operating system to run a web browser and / or the client application 4512 shown in FIG. 11. The operating system (OS) may be a Windows® OS, a Macintosh® OS, or a Linux® OS. The memory 4704 may include a combination of volatile memory (e.g., random access memory) and non-volatile memory (e.g., solid state drive and / or hard drives).Attorney Docket No.: 1532 / 3 PCT

[0098] The network interface 4710 may be a wired Ethernet interface or a Wi-Fi interface. The personal computer 4510 may be configured to access remote memory (e.g., network storage and / or cloud storage) via the network interface 4710. The UI 4708 may include a keyboard, and a pointing device (e.g., mouse). The display 4706 may be an external display (e.g., computer monitor) or internal display (e.g., laptop). In some embodiments, the personal computer 4510 may be a smart TV. In other embodiments, the display 4706 may include a holographic projector.

[0099] FIG. 14 depicts a block diagram 4800 of the mobile device 4514 of FIG. 11 in accordance with embodiments of the present disclosure. The mobile device 4514 may be any of the remote devices referenced in this disclosure. The mobile device 4514 may include an operating system to run a web browser and / or the mobile app 4516 shown in FIG. 11. The mobile device 4514 may include at least a processor 4802, a memory 4804, a UI 4806, a display 4808, WAN radios 4810, LAN radios 4812, and personal area network (PAN) radios 4814. In some embodiments the mobile device 4514 may be an iPhone® or an iPad®, using iOS® as an OS. In other embodiments, the mobile device 4514 may be a mobile terminal including Android® OS, BlackBerry® OS, Chrome® OS, Windows Phone® OS, or the like.

[0100] In some embodiments, the processor 4802 may be a mobile processor such as the Qualcomm® Snapdragon™ mobile processor. The memory 4804 may include a combination of volatile memory (e.g., random access memory) and non-volatile memory (e.g., flash memory). The memory 4804 may be partially integrated with the processor 4802. The UI 4806 and display 4808 may be integrated such as a touchpad display. The WAN radios 4810 may include 2G, 3G, 4G, and / or 5G technologies. The LAN radios 4812 may include Wi-Fi technologies such as 802.11a, 802.11b / g / n, and / or 802.1 lac circuitry. The PAN radios 4814 may include Bluetooth® technologies.

[0101] FIG. 15 depicts a block diagram 4900 of the loT device 4518 of FIG. 11 in accordance with embodiments of the present disclosure. The loT device 4518 may be any of the remote devices referenced in this disclosure. The loT device 4518 includes a processor 4902, a memory 4904, sensors 4906, servos 4908, WAN radios 4910, LAN radios 4912, and PAN radios 4914. The processor 4902, a memory 4904, WAN radios 4910, LAN radios 4912, and PAN radios 4914 may be of similar design to the processor 4802, a memory 4804, WAN radios 4810, LAN radios 4812, and PAN radios 4814 of the mobile device 4514 of FIG. 14. The sensors 4906 and servos 4908 may include any applicable components related to loTAttorney Docket No.: 1532 / 3 PCTdevices such as a monitoring device, an autonomous vehicle, a home assistant, a smart appliance, a medical device, a virtual reality device, an augmented reality device, or the like.

[0102] Any combination of one or more computer-readable medium(s) may be utilized. The computer-readable medium may be a computer readable signal medium or a computer- readable storage medium (including, but not limited to, non-transitory computer-readable storage media). A computer-readable storage medium may be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples (a non- exhaustive list) of the computer-readable storage medium would include the following: an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable readonly memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the context of this document, a computer-readable storage medium may be any tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device.

[0103] A computer-readable signal medium may include a propagated data signal with computer-readable program code embodied therein, for example, in baseband or as part of a carrier wave. Such a propagated signal may take any of a variety of forms, including, but not limited to, electro-magnetic, optical, or any suitable combination thereof. A computer-readable signal medium may be any computer-readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device.

[0104] In one embodiment, the present invention includes a cloud-based network for distributed communication via a wireless communication antenna and processing by at least one mobile communication computing device. In another embodiment of the invention, the system is a virtualized computing system capable of executing any or all aspects of software and / or application components presented herein on computing devices. In certain aspects, the computer system may be implemented using hardware or a combination of software and hardware, either in a dedicated computing device, or integrated into another entity, or distributed across multiple entities or computing devices.

[0105] By way of example, and not limitation, the computing devices are intended toAttorney Docket No.: 1532 / 3 PCTrepresent various forms of digital computers and mobile devices, such as a server, blade server, mainframe, mobile phone, personal digital assistant (PDA), smartphone, desktop computer, netbook computer, tablet computer, workstation, laptop, and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be exemplary only, and are not meant to limit implementations of the invention described and / or claimed in this document.

[0106] In one embodiment, the computing device includes components such as a processor, a system memory having a random-access memory (RAM) and a read-only memory (ROM), and a system bus that couples the memory to the processor. In another embodiment, the computing device may additionally include components such as a storage device for storing the operating system and one or more application programs, a network interface unit, and / or an input / output controller. Each of the components may be coupled to each other through at least one bus. The input / output controller may receive and process input from, or provide output to, a number of other devices, including, but not limited to, alphanumeric input devices, mice, electronic styluses, display units, touch screens, signal generation devices (e.g., speakers), or printers.

[0107] By way of example, and not limitation, the processor may be a general-purpose microprocessor (e.g., a central processing unit (CPU)), a graphics processing unit (GPU), a microcontroller, a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field Programmable Gate Array (FPGA), a Programmable Logic Device (PLD), a controller, a state machine, gated or transistor logic, discrete hardware components, or any other suitable entity or combinations thereof that can perform calculations, process instructions for execution, and / or other manipulations of information.

[0108] In another embodiment, multiple processors and / or multiple buses may be used, as appropriate, along with multiple memories of multiple types (e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core).

[0109] Also, multiple computing devices may be connected, with each device providing portions of the necessary operations (e.g., a server bank, a group of blade servers, or a multiprocessor system). Alternatively, some steps or methods may be performed by circuitry that is specific to a given function. According to various embodiments, the computer system may operate in a networked environment using logical connections to local and / or remoteAttorney Docket No.: 1532 / 3 PCTcomputing devices through a network. A computing device may connect to a network through a network interface unit connected to a bus. Computing devices may communicate communication media through wired networks, direct-wired connections or wirelessly, such as acoustic, RF, or infrared, through an antenna in communication with the network antenna and the network interface unit, which may include digital signal processing circuitry when necessary. The network interface unit may provide for communications under various modes or protocols.

[0110] In one or more exemplary aspects, the instructions may be implemented in hardware, software, firmware, or any combinations thereof. A computer readable medium may provide volatile or non-volatile storage for one or more sets of instructions, such as operating systems, data structures, program modules, applications, or other data embodying any one or more of the methodologies or functions described herein. The computer readable medium may include the memory, the processor, and / or the storage media and may be a single medium or multiple media (e.g., a centralized or distributed computer system) that stores the one or more sets of instructions. Non-transitory computer readable media includes all computer readable media, with the sole exception being a transitory, propagating signal per se. The instructions may further be transmitted or received over the network via the network interface unit as communication media, which may include a modulated data signal such as a carrier wave or other transport mechanism and includes any delivery media. The term “modulated data signal” means a signal that has one or more of its characteristics changed or set in a manner as to encode information in the signal.

[0111] Storage devices and memory include, but are not limited to, volatile and nonvolatile media such as cache, RAM, ROM, EPROM, EEPROM, FLASH memory, or other solid state memory technology; discs (e.g., digital versatile discs (DVD), HD-DVD, BLU- RAY, compact disc (CD), or CD-ROM) or other optical storage; magnetic cassettes, magnetic tape, magnetic disk storage, floppy disks, or other magnetic storage devices; or any other medium that can be used to store the computer readable instructions and which can be accessed by the computer system.

[0112] The various illustrative logical blocks, modules, elements, circuits, and algorithms described in connection with the embodiments disclosed herein may be implemented as electronic hardware, computer software, or combinations of both. To clearly illustrate this interchangeability of hardware and software, various illustrative components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. WhetherAttorney Docket No.: 1532 / 3 PCTsuch functionality is implemented as hardware or software depends upon the particular application and design constraints imposed on the overall system. Skilled artisans may implement the described functionality in varying ways for each particular application (e.g., arranged in a different order or partitioned in a different way), but such implementation decisions should not be interpreted as causing a departure from the scope of the present invention.

[0113] The flowchart and block diagrams in the Figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of code, which comprises one or more executable instructions for implementing the specified logical function(s). It should also be noted, in some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and computer instructions.

[0114] The computer program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other devices to cause a series of operational steps to be performed on the computer, other programmable apparatus or other devices to produce a computer implemented process such that the instructions which execute on the computer or other programmable apparatus provide processes for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0115] The corresponding structures, materials, acts, and equivalents of all means or step plus function elements in the claims below are intended to include any structure, material, or act for performing the function in combination with other claimed elements as specifically claimed. The description of the present invention has been presented for purposes of illustration and description but is not intended to be exhaustive or limited to the invention in the form disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the invention. The embodiment was chosenAttorney Docket No.: 1532 / 3 PCTand described in order to best explain the principles of the invention and the practical application, and to enable others of ordinary skill in the art to understand the invention for various embodiments with various modifications as are suited to the particular use contemplated.

[0116] The descriptions of the various embodiments of the present invention have been presented for purposes of illustration but are not intended to be exhaustive or limited to the embodiments disclosed. Many modifications and variations will be apparent to those of ordinary skill in the art without departing from the scope and spirit of the described embodiments. The terminology used herein was chosen to best explain the principles of the embodiments, the practical application or technical improvement over technologies found in the marketplace, or to enable others of ordinary skill in the art to understand the embodiments disclosed herein.

Claims

Attorney Docket No.: 1532 / 3 PCTCLAIMSWhat is claimed is:

1. An artificial intelligence based voice evaluation system for assisting diagnosis and treatment of voice disorders, the artificial intelligence based voice evaluation system comprising:at least one sensor;at least one remote device including at least one user interface;at least one remote server including at least one database;at least one software platform including an analytics engine and a calibration engine;andwherein the at least one sensor is operable to capture vocal data, wherein vocal data includes vocal health data and vocal function data;wherein the at least one sensor, the at least one remote device, the at least one remote server, and the at least one software platform are in network communication;wherein the at least one sensor is operable to transmit the captured vocal data to the at least one remote device, the at least remote server, and the at least one software platform;wherein the at least one remote device is operable to display the captured vocal data via the at least one user interface; wherein the at least one remote server is operable to store the captured vocal data in the at least one database; andwherein the analytics engine is operable to determine a vocal health of a user based on the vocal data.Attorney Docket No.: 1532 / 3 PCT2. The artificial intelligence based voice evaluation system of claim 1 , wherein the analytics engine further includes a machine learning component.

3. The artificial intelligence based voice evaluation system of claim 2, wherein the vocal data includes a multi-modal dataset, wherein the multi-modal dataset includes aerodynamic data and acoustic data.

4. The artificial intelligence based voice evaluation system of claim 3, wherein the analytics engine is further operable to generate at least one recommendation based on the determined user vocal health.

5. The artificial intelligence based voice evaluation system of claim 4, wherein the at least one recommendation includes a treatment recommendation, a rehabilitation recommendation, or a physician recommendation.

6. The artificial intelligence based voice evaluation system of claim 5, wherein the at least one sensor includes an audio sensor, an electroglottograph sensor, a humidity sensor, an imaging sensor, a movement sensor, an aerodynamic sensor, and / or a temperature sensor.

7. The artificial intelligence based voice evaluation system of claim 6, wherein the calibration engine is operable to generate at least one sensor recommendation based on the vocal data, wherein the at least one sensor recommendation includes at least one of positioning, digital gain control, and / or signal filtering.

8. The artificial intelligence based voice evaluation system of claim 1, wherein the vocal data includes vocal imaging data, vocal aerodynamic data, vocal acoustic data, and self-evaluation data.

9. An artificial intelligence based voice evaluation system for assisting diagnosis and treatment of voice disorders, the artificial intelligence based voice evaluation system comprising:at least one vocal evaluation device including at least one sensor;Attorney Docket No.: 1532 / 3 PCTat least one user interface; andan analytics engine;wherein the at least one sensor is operable to capture vocal data, wherein vocal data includes vocal health data and vocal function data;wherein the at least one vocal evaluation device is operable to display the captured vocal data via the at least one user interface; wherein the at least one vocal evaluation device is operable to store the captured vocal data; andwherein the analytics engine is operable to determine a vocal health of a user based on the captured vocal data.

10. The artificial intelligence based voice evaluation system of claim 9, wherein the vocal data includes a multi-modal dataset, wherein the multi-modal dataset includes aerodynamic data and acoustic data.

11. The artificial intelligence based voice evaluation system of claim 10, wherein the analytics engine is further operable to generate at least one recommendation based on the determined user vocal health.

12. The artificial intelligence based voice evaluation system of claim 11 , wherein the at least one recommendation includes a treatment recommendation, a rehabilitation recommendation, or a physician recommendation.

13. The artificial intelligence based voice evaluation system of claim 12, wherein the at least one sensor includes an audio sensor, an electroglottograph sensor, a humidity sensor, an imaging sensor, a movement sensor, an aerodynamic sensor, and / or a temperature sensor.

14. The artificial intelligence based voice evaluation system of claim 13 further comprising a calibration engine, wherein the calibration engine is operable to generate at least one sensorAttorney Docket No.: 1532 / 3 PCTrecommendation based on the vocal data, wherein the at least one sensor recommendation includes at least one of positioning, digital gain control, and / or signal filtering.

15. The artificial intelligence based voice evaluation system of claim 9, wherein the vocal data includes vocal imaging data, vocal aerodynamic data, vocal acoustic data, and selfevaluation data.