Vowel-centered speech modification
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2026-02-06
- Publication Date
- 2026-08-13
Smart Images

Figure US2026014371_13082026_PF_FP_ABST
Abstract
Description
Attorney Ref. No. 1033.1004. WOPCT PATENTVOWEL-CENTERED SPEECH MODIFICATIONCROSS-REFERENCE TO RELATED APPLICATIONS
[0001] This application claims priority and benefit from the U. S. Provisional Patent Application 63 / 755.979, filed February 7, 2025, and titled, “VOWEL-CENTERED SPEECH MODIFICATION,” and U. S. Provisional Patent Application 63 / 903,553, filed October 22, 2025, and titled, “VOWEL-CENTERED ACCENT MODIFICATION IN PILOTS,” which are incorporated herein by reference in their entirety for all purposes.BACKGROUND
[0002] Language is the very thread that connects human beings. Without language, one would have no way to communicate wants, needs, indicate help, or express complex ideas and theories. Housed within language is the domain of intelligibility or being understood by others when speaking. One can have superb language skills, but the same person may struggle to be understood by others due to underdeveloped speech skills that can result in negative consequences presented by limits on relationships and communication, such as personal and work opportunities. For example, the aviation industry requires pilots to speak with a high level of intelligibility so they are understood across a radio intercom when operating an aircraft. The U. S. Federal Aviation Administration (FAA) closely follows the International Civil Aviation Organizations (ICAO) English Proficiency Descriptors to set pilot standards. These ICAO requirements were created secondary to support strong safety communication standards in the aviation industry after the repeated tragic loss of life over what was attributed to either air traffic controllers'’ (ATC) inability to understand the pilot’s English speech or the pilot’s failure to understand the English used by ATCs. All these tragic events were preventable with improved pilot intelligibility and understanding. The ICAO, FAA, and FAA counterparts in all countries across the globe adopted English as their common language. These regulatory organizations require all pilots from all countries to be proficient in English to receive their pilot’s license.
[0003] Other industries face similar challenges with speech intelligibility, such as professionals in healthcare, aerospace, maritime, miliary, engineering, science, technology, technological support, clergy, business, customer service, and the like. The common thread is an industry thatAttorney Ref. No. 1033.1004. WOPCT PATENTadopts a language as ubiquitous industry-wide or globally or for a particular use case such as workers interacting with customers in a foreign country in a language not native to the workers and native to the customers in the foreign country. According to the U. S. Census Bureau's American Community Survey, approximately 68 million people in the United States spoke a language other than English at home (2019). Additionally, approximately one in four or 26% of physicians working in the United States are born in other countries and 30% of practicing U. S. Catholic clergy are born in other countries. Oftentimes, businesses provide outsourced customer service and / or technical support from call centers located outside the U. S. Each of these groups of non-native English speakers, regardless of their physical location, are expected to communicate in English as they engage with customers.
[0004] In aviation, for example, when a prospective pilot’s native language (LI) is not English, they may have accents that interfere with their speech intelligibility. Accents consist of traits carried over from a native language to the learned foreign language, which is English in this example. The speech -language pathologist (SLP) is the ideal professional to serve those working on improving intelligibility through accent modification and reduction treatment sendees. Accent modification and reduction treatment services consist of targeting the specific ways the individual’s speech is audibly different from the language they speak as a foreign language. Regardless of the speaker’s profession, clear communication through speech is a foundational skill required to provide professional services globally. Individuals from different countries communicate multi -lingually for personal and professional reasons. Verbal communication consists of three principal components: the sender, the message provided by the sender, and the receiver. Clear verbal communication is the sought-after transmission of speech, which must be intelligible to a receiver for the sender- speaker to be successful. Vowel errors in speech are the most difficult sounds to correct and modify and often form foundational issues with a speaker’s intelligibility, in particular with target phraseology in industry specific speech, such as aviation terminology or other terminology from other industries. Most non-native English speakers, for example, struggle with one or more vowel sound errors that makes their intelligibility lower, sometimes below thresholds for professional licensing and effective, safe communication in their personal and professional lives.
[0005] However, conventional speech therapy focuses on consonant based error correction. SLPs help speakers correct consonant sounds with clear instructions to move their tongue, lips,Attorney Ref. No. 1033.1004. WOPCT PATENTor mouth in a different way or shape. There are “touch points” of the tongues, lips, or mouth when forming consonant sounds so corrective behavior can be explained by the SLPs using specific instructions to change the manner in which the speaker is forming the words with their tongue, lips, or mouth. Unlike consonants, vowels do not have corresponding articulatory contact points of the tongue, lips, or other oral structures. Instead, vowel production requires the speaker to alter the shape and position of the vocal tract which includes the tongue, lips, oral cavity, and pharynx in order to modify the resonant properties of airflow from the lungs. It is the vowel sounds that most significantly influence perceived accent and intelligibility, yet they are rarely targeted due to their extreme difficulty to modify, change, or improve." Explaining how¬ to change a vowel sound to a speaker is difficult, abstract, and often fails in treatment so SLPs focus their work with speakers on improving their consonant sounds. Changing vowel sounds helps improve the speakers intelligibility vastly more than changing the consonant sounds. However, no conventional treatment exists for helping speakers improve their vowel sounds that is effective. Further, no conventional treatment focuses on industry specific speech modification for sounds, words, phrases, sentences, and conversational discourse that are critical to allow a speaker to highly function in an industry, profession, or social environment. Those speakers that wish to gain employment or integrate within a community that speaks or has adopted a language, such as English, other than the speakers native language will lose professional and personal opportunities if they are not proficient in that adopted primary language.
[0006] What is needed is a reliable, effective speech modification system and method that helps speakers modify their speech in desired ways for industry or professional and social environment specific sounds, words, phrases, sentences, and conversational discourse.BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Non-limiting and non-exhaustive embodiments of the invention are described with reference to the following drawings. In the drawings, like reference numerals refer to like parts throughout the various figures, unless otherwise specified, wherein:
[0008] FIG. 1A is an example method of vowel-centered speech modification techniques according to aspects of this disclosure.
[0009] FIG. IB is another example method of vowel-centered speech modification techniques disclosed in this application.Attorney Ref. No. 1033.1004. WOPCT PATENT
[0010] FIG. 1C is still another example method of vowel -centered speech modification techniques according to this disclosure.
[0011] FIG. 2 is an example system diagram of vowel-centered speech modification techniques disclosed herein.
[0012] FIG. 3A is an example speech evaluation passage.
[0013] FIG. 3B is another example speech evaluation passage.
[0014] FIG. 4 is a vowel quadrilateral related to aspect of this disclosure.
[0015] FIGS. 5A and 5B are example spectrograms showing formant 1 (Fl) and formant 2 (F2) in each example.
[0016] FIG. 5 C is an example method of providing a user with visual output based on user input of their speech without a spectrogram in accordance with aspects of this disclosure.
[0017] FIG. 6 is a chart of Fl and F2 measurements in Hertz (Hz) of vowel sounds in American Standard English.
[0018] FIGS. 7A and 7B is a chart of Aviation Vocabulary Categories to create an individualized vowel-centered speech modification program according to aspects of this disclosure.
[0019] FIG. 8 is an example visual output in accordance with the vowel-centered speech modification techniques disclosed herein.
[0020] FIG. 9 is another example visual output according to aspect of the vowel-centered speech modification techniques disclosed herein.DETAILED DESCRIPTION
[0021] The subject matter of embodiments disclosed herein is described here with specificity to meet statutory requirements, but this description is not necessarily intended to limit the scope of the claims. The claimed subject matter may be embodied in other ways, may include different elements or steps, and may be used in conjunction with other existing or future technologies. This description should not be interpreted as implying any particular order or arrangementAttorney Ref. No. 1033.1004. WOPCT PATENTamong or between various steps or elements except when the order of individual steps or arrangement of elements is explicitly described.
[0022] Embodiments will be described more fully hereinafter with reference to the accompanying drawings, which form a part hereof, and which show, by way of illustration, exemplary embodiments by which the systems and methods described herein may be practiced. The systems and methods may. however, be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will satisfy the statutory requirements and convey the scope of the subject matter to those skilled in the art.
[0023] Vowel-centered accent modification with visual feedback has demonstrated effectiveness in reducing accented language skills in foreign speakers, and specifically in specialized professions where language proficiency is essential to obtaining licensure and ensuring critical safety standards are met. For example, the universally adopted language in the global aviation industry is English. Pilots from across the globe are trained to operate aircraft with only a fraction of them being native English speakers. These pilots that have a native language (LI) other than English often struggle to be understood well in their critical functions of operating their aircraft due to accents they apply to their English speech. Unfortunately, some of these accents of non-native English speakers negatively impact their intelligibility of critical safety words and functions while operating their aircraft. This phenomenon leads to serious, and sometimes critical, safety events that impacts the life safety of the pilots, passengers, and bystanders on the ground over which the aircraft are flying. Licensure of pilots requires a minimum intelligibility in English with accent modification often required for pilots to meet licensing requirements and avoid these critical safety events.
[0024] The disclosed systems and methods provide vowel-centric accent modification techniques that are incorporated into a software, or other computer-readable media, diagnostic and therapeutic techniques, hybrid models, training, and other therapy techniques employed to improve a user’s industry specific speech. Industry specific speech are those sounds, words, phrases, sentences, and conversational discourse that impact the job function of the user. With pilots, for example, speech relating to operating aircraft and communicating with air traffic controllers and other pilots are critical functions that pose extreme safety risks to all pilots, passengers, and bystanders near the aircraft. The same or similar techniques can apply to otherAttorney Ref. No. 1033.1004. WOPCT PATENTprofessions, such as healthcare, business, customer service, legal, and the like to improve intelligibility of a user within their specific professional industry. The examples disclosed herein focus on the aviation industry but can be employed by any other profession, industry, or likeminded group with specific sounds, words, phrases, sentences, and / or conversational discourse that relate to their industry or group.
[0025] For example, the disclosed systems and methods can be incorporated into an online subscription service, application, training manual, or any combination thereof. Any business that needs to reduce accentedness in employees, students, or other stakeholders, such as international aviation academies, government entities, such as the Federal Aviation Administration (FAA), international engineering groups, or hospitals that employ international locum physicians, could implement the disclosed techniques, systems, and methods to meet international and profession language proficiency standards, such as proficiency in English. Most professions do not conventionally incorporate any speech therapy based training. Generalized speech therapy most often focuses on consonant-centered speech therapy techniques because of the ability to describe the placement of the user’s tongue and lips and the shape of their mouth when forming sounds. The disclosed vowel-centric speech therapy techniques described herein improve delivery of language training across various disciplines for many reasons.
[0026] FAA-funded research indicates that accent, speech rate, and pronunciation adversely affect pilots' ability to understand word meanings to a greater extent than radio technique or equipment quality (Prinzo et al., 2010). The International Civil Aviation Organization (ICAO) has recognized communication deficiencies as contributing factors in fatal aviation accidents and, since 2008, has required all pilots and air traffic controllers to demonstrate English language proficiency. These accidents are preventable with help from the disclosed vowel-centered accent modification techniques. To obtain a pilot's license, designated pilot examiners assess a pilot's English pronunciation ability as part of the flight check for non-native English-speaking pilots. Modern-day aviation is an international industry, yet accented aviation English remains a safety concern acknowledged by global regulatory bodies and industry professionals. The disclosed vowel-centric speech modification techniques can play a vital role in reducing accentedness in pilot candidates and licensed pilots, thereby improving overall intelligibility and increasing aviation safety margins.Attorney Ref. No. 1033.1004. WOPCT PATENT
[0027] To reduce accentedness and increase aviation industry safety, the disclosed techniques use vowel-centered speech modification and visual feedback with speech therapy and treatment for pilots. Further, the disclosed techniques adapts oronym complexity for the user-specific words and program sessions based on the user’s performance data along with LI specific pronunciations and industry specific word choice. An oronym, as used herein, is a sequence of words or syllables that, when spoken using the phonological patterns familiar to the user, produces acoustic output that approximates the target pronunciation. The oronym may originate from the user's first language (LI), the target second language (L2), or a combination thereof. The oronym leverages the user's existing articulatory habits to produce target sounds without requiring explicit articulatory instruction. For example, a Spanish-speaking pilot learning to produce the English numeral 'one' / wAn / , a mandatory ICAO standard pronunciation, may use the familiar Spanish name 'Juan' as an oronym, whose natural articulation produces vowel formant values (Fl and F2) within the acceptable range of the target English vowel. The user produces the target sound through familiar motor patterns rather than conscious manipulation of articulatory placement. Conventional accent modification techniques rely on consonant-centric speech modification and do not utilize forced prosodic alignment to teach the user intelligibility. These traditional approaches fail to predict speech modification needs related to the user’s LI and do not adapt to the user’s performance as they progress through a program. Further, they focus on consonant pronunciation modification rather than vowel pronunciation because modification of a user’s consonant pronunciation can easily be described orally or through visual example using cues and instructions regarding shape and placement of the user’s tongue, lips, and mouth. Vowel pronunciations require more complex and confusing cues and instructions to users, and as such, have a lower success rate to correct than consonants do. This disclosure explains how to modify a user’s accentedness quickly and efficiently by coupling vowel-centric pronunciation with individualized oronyms and user education of the Fl and F2 sounds with real-time visual cues to the user to make vowel-centric pronunciation adjustments and ultimately speech modification to reduce accentedness.
[0028] The examples described herein focus on the aviation industry, but they can be modified to improve accentedness for any user in any other industry, profession, or social environment. Here, these examples are described in relation to a user of a program, such as an algorithm that controls a “user’s” interaction with an application on their electronic device, such as a mobileAttorney Ref. No. 1033.1004. WOPCT PATENTphone or laptop computer. However, these same or similar techniques can be employed with a SLP therapist present to help guide a “user.” User, in the context of this disclosure, can be substituted for “speaker” and means the person, client, patient, speaker, user, or other human being undergoing the disclosed treatment technique to modify their speech through the vowelcentric techniques described here.
[0029] Regardless of the environment or use case, the disclosed techniques first evaluate a user’s speech using a standardized reading passage read by the user. The reading passage is vowel-centric in this disclosure because the treatment techniques are vowel-centric. The reading passage can be a speech evaluation passage typical in speech therapy for users speaking Standard American English to help SLPs diagnose speakers (in this disclosure “users”) of various speech disorders and accentedness. The vowel-centric reading passage is a standardized passage comprising substantially all Standard American English phonemes, with particular emphasis on vowel sounds critical to intelligibility assessment. The passage is configured to assess a user's articulation, intelligibility, fluency, voice quality, and prosody. By concentrating diagnostic content on vowel production, the passage enables targeted identification of accentedness patterns that most significantly affect listener comprehension. A vowel-centric reading passage can be used in this disclosure to adjust the disclosed comprehensive approach to evaluating speech to focus more on vowels compared to traditional speech passages used to evaluate users. The vowel-centric nature of the reading passage allow the disclosed techniques to readily identify vowel-pronunciation deficiencies of its users. Improvement of these vowel deficiencies has a greater impact on correcting accentedness than improving consonant deficiencies. Standard American English comprises approximately 18 vowel sounds, including monophthongs, diphthongs, and rhotic vowels. As used herein, the term "vowels" refers collectively to all monophthong, diphthong, and rhotic vowel sounds unless otherwise specified.
[0030] The vowel-centric reading passage that focuses on vowel pronunciation, represents each vowel in Standard American English equally or in a targeted ratio of vowels to consonants. It also includes vowels specifically selected to determine the user’s vowel pronunciations deficiencies that lead to their accentedness and directly impact their intelligibility due to their native language speech trait transfers. In some examples, the vowel-centric reading passage is specific to terms frequently spoken in an industry, profession, or social situation. The frequency of the vowel sounds in the core or critical terms frequency spoken in the industry, profession, orAttorney Ref. No. 1033.1004. WOPCT PATENTsocial situation can vary depending on the words. Each vowel is presented in the reading passage at least four times. Vowels that are 50% accurate or less are potential targets. Further, in some industries, such as aviation, the vowel sounds / a / , / i / , / e / , / i / , / ei / , and / ɑ / are spoken at a higher frequency because the industry regulated terms more often require use of those vowel sounds. The aviation industry specific reading passage includes, for example, each vowel sound presented a sufficient number of times to identify error patterns, wherein vowel sounds produced incorrectly at or above a predetermined threshold are prioritized against the frequency of occurrence of those vowel sounds in industry-regulated terminology, to target the most often used sounds for the user's guided speech work in the user sessions. This technique to weight the target vowel work to the sounds most frequently occurring in the user's industry leads to the most significant improvement in the user's intelligibility in the most efficient amount of time.
[0031] In some embodiments, the vowel-centric speech modification program includes a speech sample for the users at the beginning of the evaluation or a user session. The speech sample includes a prompt that the user is asked to answer or discuss in open-ended speech. The user’s speech sample is used to validate the guided speech intervention for that user session. The vowel-centric speech modification program can request or require a speech sample at each user session or periodically through a multi-session program either at set intervals or when manually requested by the user or another person or entity, such as an SLP. The vowel-centric speech modification program analyzes the speech sample provided by the user to determine vowel errors that already exist or are new areas of targeted intervention for a particular user session. For example, after the initial user evaluation, which may include a speech sample, and creation of multiple sessions for the user’s individualized vowel-centric speech modification program, the user may be required or requested to provide another speech sample after completing 5 user sessions to ensure that the subsequent user sessions continue to meet the user at their level of individual need. Any suitable frequency of requiring or requesting the user to provide a speech sample can be used, including at the beginning of every session, if desired. Alternatively, the speech samples are not required for the vowel-centric speech modification program to benefit the user. The speech sample supplements the initial evaluation of the user by the vowel-centric reading passage using the industry, profession, or situation specification reading passage.
[0032] Once the users are evaluated, the disclosed vowel-centric speech modification techniques create a vowel-centric speech modification program that adaptively progresses a userAttorney Ref. No. 1033.1004. WOPCT PATENTto improve their speech. This vowel-centric speech modification program has varying levels of speech hierarchy that starts with user sessions that isolate vowel sounds and progresses to user sessions that combine consonant and vowel sounds. Users can start with the vowel-centric speech modification program at any level, such as the word level if that is appropriate for their level of speech. More advanced users can enter the program at an even higher speech hierarchy level, as needed. Once the user is proficient at pronouncing combined consonant and vowel sounds, the vowel-centric speech modification program progresses to mono- and bisyllabic words, phrases, sentences then semi- structured speech and / or conversational discourse. As the users perform in each session, the disclosed technique adaptively adjusts the next session based on the user’s performance in either the last session or a cumulation of all or a portion of the previous sessions. The users are given gamified and industry specific visual cues that indicate how the user must change the placement or shape of their tongues, lips, and mouth to improve a spoken vowel sound. Vowel sounds are made with an open vocal tract, which is nearly impossible to modify without visual feedback. Because correct vowel sounds are the foundation of intelligibility, the visual feedback relates to the changes the user needs to make to improve their vowel pronunciation. Other types of user feedback, such as auditory feedback, have failed to change a user’s vowel pronunciation and intelligibility. Visual feedback allows the users to receive treatment with an SLP and / or through applications on electronic devices, such as mobile phones and computers, which allows for effective and scalable treatment for the users at their leisure. Such easily accessible accent modification treatment that employs the disclosed techniques increases user adherence to and completion of the treatment.
[0033] Further, the disclosed techniques select industry or profession specific sounds, words, phrases, sentences, and conversational discourse for the user sessions of the vowel-centric speech modification program. These selected sounds, words, phrases, sentences, and conversational discourse topics can also be specific to social environments or other specific use cases. For example, the sounds, words, phrases, sentences, and conversational discourse for the user sessions can be selected for the aviation industry that focus on aviation communications between pilots and air traffic control and support new pilots is reaching proficiency in speaking Standard American English to pass the oral exam of the licensing requirements.
[0034] Turning now to FIGS. 1A - 1C, the disclosed methods 100A, 100B, 100C and systems evaluate a user’s phonological challenges that is used to create an individualized speechAttorney Ref. No. 1033.1004. WOPCT PATENTmodification program that uses visual cues to adaptively improve the user’s speech over the course of multiple program sessions. During the evaluation process 100A, the user is prompted to read a reading passage aloud that is recorded and evaluated. Optionally, the vowel-centric speech modification program can require or request a user submits a 30-second (or other time interval) open-ended speech sample at the start of a user session that is recorded and analyzed. In an example, the reading passage is a vowel-centric reading passage specific to an industry, profession, social situation, or other specific use case, which emphasizes the vowels of a target language. Here in this disclosure, the examples are focused on users reading the vowel-centric reading passage in Standard American English for purposes of earning pilot licensure. English is the universal language in the global aviation industry and proficiency in speaking it is required for licensing all pilots across the globe as required by ICAO, a subgroup of United Nations. In other examples or use cases, the reading passage and optionally the speech sample can be a different language.
[0035] The vowel-centric speech modification program receives the user’s reading passage 102 and optionally receives the user’s speech sample then extracts frequency values for a first formant (Fl) and a second formant (F2). Fl and F2 are measured in frequency values, such as Hertz (Hz). Fl and F2 are two low fundamental or resonant frequencies of a spoken sound and are primarily used to identify and distinguish vowel sounds. The vowel-centric speech modification program identifies the Fl and F2 by extracting frequency data from the input speech data of the user. The input speech data from the user can be either from the initial evaluation using the industry specific reading passage or from a user speech sample in a particular user session. The input speech data can also be analyzed in real-time or near real-time (collectively referred to hereafter as “real-time”) as the user speaks using the individualized oronyms and selected sounds and words prepared for them in a user session. The real-time analysis of the user’s speech translates to continuous Fl and F2 data that tracks the user’s realtime speech, which can be used to adjust visual output so the user can see how to adjust their pronunciation in real-time.
[0036] Alternatively, the vowel-centric speech modification system can create a spectrogram of the vowel-centric reading passage sample 104 or the user’s real-time speech. A spectrogram shows pictures and patterns of the user’s speech. Translating spectrograms into digestible instruction is difficult for a user or anyone untrained in spectrographic interpretation. In someAttorney Ref. No. 1033.1004. WOPCT PATENTembodiments, the vowel-centric speech modification system creates a spectrogram and determines a first formant (Fl) and a second formant (F2) from the spectrogram 106. The Fl and F2 in a spectrogram are resonant frequencies of the human vocal tract that appear as dark bands of concentrated acoustic energy on a spectrogram
[0037] The Fl and F2 frequency values are measured against respective Fl intelligibility criterion and F2 intelligibility criterion regardless of whether they are extracted from the input data in real-time or through a spectrogram. In some examples, the system both extracts the data in real-time and creates a spectrogram to output. A spectrogram can be useful to a specialist trained in spectrogram interpretation that further helps support or analyze the progress of the user. For example, the Fl and F2 intelligibility criterion are a minimum standard, threshold, or other criterion that makes for intelligible speech and is compared to the frequency values of the user’s pronunciation while speaking a certain vowel in the reading passage. If the Fl and F2 frequency values are below, above, or outside a threshold frequency for an intelligibility criterion, then the method determines the vowel is a target intelligibility vowel for the user to improve 108. A user with spoken vowels outside the expected frequency range speaks with an accent and their intelligibility decreases the farther outside the frequency range their spoken vowels. To reduce or eliminate the accent, the user has to change the size of their mouth or throat with their tongue for their spoken vowels to fall within the expected frequency range or pattern of a native speaker of the target language. The method then individually selects target words, oronyms. phrases, phonetic sequences, or the like from the identified target intelligibility vowels 110. Each of the selected target words, oronyms, phrases, phonetic sequences or the like relate to the user’s speech intelligibility in the target language. The method creates an individualized user vowel-centric speech modification program that has multiple user sessions 112. The program is based on the individually selected target vowels and progresses the user through a series of multiple sessions that adapt to the user’s speech performance.
[0038] FIG. IB shows a method that has a user engaging in a session of the speech modification program 100B. This user session 100B shown in FIG. IB can be based on the user’s evaluation, such as the user’s performance of a reading passage as described in FIG. 1A. The user session shown in FIG. IB adaptively progresses a user through multiple sessions based on the user’s performance in any single or multiple sessions. The user begins by starting a session with selected target words, such as those described in FIG. 1A. The method receives the input as theAttorney Ref. No. 1033.1004. WOPCT PATENTuser’s speech from reading the target words to engage with a user session of the vowel-centric speech modification methods and systems 114 described herein. The method then extracts or identifies the Fl and F2 frequency values continuously throughout the user session or a portion thereof. The method creates visual output for the user based on the Fl and F2 frequency values 118. The visual output has a first visual cue related to the Fl and a second visual cue related to the F2. The first visual cue and the second visual cue can be vertical or horizontal dimensions placed on a horizontal viewpoint of the visual output. To be understood and maintain high quality speech intelligibility, users must have their Fl and F2 frequency vowel values in their target words spoken to be maintained in the visual output within the first visual cue and second visual cue boundaries. To maintain speech intelligibility, the user's Fl and F2 frequency values for each target word must remain within the boundaries defined by the first visual cue and second visual cue. Within the stream of intelligibility, there can be multiple visual cue levels that correlate to multiple proficiency levels. In an example, there is a tunnel with an outer portion of the tube representing proficient pronunciation of the target words and an inner portion of the tube representing mastery of pronouncing the target words.
[0039] The first visual cue and the second visual cue dynamically change based on the Fl and F2 frequency values of the user's real-time speech 120 as the user progresses through a user session of target words. The visual cues adapt to the user's real-time speech and visually indicate the changes the user needs to make to correct any vowel (or consonant) pronunciations. The visual cues correlate to changes in the size of the user's mouth or throat with their tongue for their spoken vowels to be intelligible and within the expected range. This means that the first visual cue and / or the second visual cue visually indicate to the user how to adjust the placement, shape, or size, of the user's tongue, lips, or mouth. In an aviation example, the entire aircraft drops below the Fl frequency value, which indicates the user needs to close the mouth more and raise the tongue. In another aviation example, the aircraft nose shifts beyond the F2 frequency value, which indicates the user needs to change the front to back position of their tongue. A user learns best and most efficiently when they are able to correct in a guided corrective learning environment. The visual output helps them stay within the zone of guided corrective feedback so they develop only good speech skills rather than practice incorrect pronunciation throughout user sessions.Attorney Ref. No. 1033.1004. WOPCT PATENT
[0040] FIG. 1C shows an example method of selecting individualized targets words and / or oronyms for an individual user 100C. Tier 1 includes selecting industry, profession, or social environment specific oronyms or words to begin a progressive speech modification program 126. These industry / profession / social environment specific terms relate to a group of sounds, words, phrases, or conversational discourse topics that relate to the specific target industry, profession, or social environment. In some examples like the aviation industry, target words include various categories of words and phrases that allow pilots to perform core functions of their job. In the aviation industry for example, safety critical words, aircraft operations, crew communications, and numbers and altitudes are a few of many categories of words that pilots must clearly communicate and comprehend to successfully perform their job. Other industries or professions and different social situations have different categories of words or oronyms specific to their specific use case. Each word or phrase in the industry / profession / social environment specific environment has a phonetic pronunciation that enables comprehensive LI specific target language training for the speakers.
[0041] Tier 2 of selecting individualized target words or oronyms for a user 100C includes adding known native language transfer patterns of a user 128. Native language transfer patterns can be based on a user’s profile with a known or documented transfer pattern and / or native language predictive transfer patterns based on typical users that speak the same native language as the user. Transfer patterns are rules, structures, speech sounds, and vocabulary from the user’s first language that impact their ability to learn, speak, and comprehend a non-native language. Transfer patterns can be both negative and positive. Negative transfer patterns are known or identified difficulties in a user’s ability to pronounce a letter or sound or communicate a concept that may not exist or may be different in the user’s native language. Positive transfer patterns are known of identified similarities that allow a user to transfer basic linguistic understanding from their native language to a non-native language. The Tier 2 goal is to progress the user’s pronunciation based on their native language transfer patterns, both positive and negative.
[0042] Tier 3 of selecting individualized target words or oronyms for a user 100C include selecting oronyms to advance the user’s progressive pronunciation guidance based on a user’s individual performance in an evaluation or a user session of the vowel-centric speech modification program 130. For example, a user reads the vowel-centric reading passage. TheAttorney Ref. No. 1033.1004. WOPCT PATENTfirst session of the created vowel-centric speech modification program focuses on the phonological challenges noted by the disclosed methods and systems based on the user’s reading passage evaluation. All subsequent user sessions are evaluated and adjusted, if needed, based on the user’s progress during each successive user session. The work in each successive user session adapts to the user’s current level of speech rather than progressing without checking the user’s speech progress. Such adaptive progression through a treatment plan is more akin to in-person therapy with an SLP rather than a user engaging with an application or computing device. Adaptive progression also provides the gold standard care for improve speech in general, and specifically in improving accentedness in users learning a new language.
[0043] In an example, the words or oronyms selected in Tier 3 become more advanced as the user performs better, stay at a similar difficulty level if the user performs at about the same level, and may become simpler if the user regresses. In the disclosed vowel-centric speech modification methods and systems, the user typically progresses through learning acceptable sounds of each vowel; then learning sounds of each vowel plus a consonant: then learning consonant sounds plus each vowel; then to words, phrases, semi-structured speech, and conversational discourse. A user can enter the program at any level based on their performance on the reading passage and / or speech sample. Regardless of the level of entry of the user, the user progresses, adaptively through the program based on their verbal performance. The performance can be the immediately preceding performance, trends on multiple preceding performances, and any other performance evaluation methods. For example, the disclosed methods and systems can periodically re-evaluate the user to ensure the user is learning at the most appropriate level for the user. The disclosed methods and systems can also progress a user based on the immediately preceding performance only or a cumulative analysis of all or a portion of immediately preceding performance. If the disclosed methods and systems progress the user based on a portion of the preceding performances, they can be limited to a capped number of preceding user performances, such as the preceding 10 user performances, and / or to all user performances within a preceding time period, such as a week for a frequent user or a month for a less frequent user.
[0044] Tiers 1-3 work together to layer on each other to progressively customize the speech modification treatment the user receives. Each of the Tiers 1-3 work together to provide the most individualized vowel-centric speech modification for the user. All the words and phrasesAttorney Ref. No. 1033.1004. WOPCT PATENTthe disclosed methods and systems give to the user are selected to be industry or profession specific, as disclosed in the Tier 1 selections. The Tier 2 selected words and phrases are layered upon the Tier 1 industry or profession specific choices to always account for the user’s native language transfer patterns from either or both of their user profile (known transfer patterns specific to the user) or from empirical data gathered about typical transfer patterns of others that speak their native language. Once the user completes the evaluation and for all subsequent user sessions, the Tier 3 selected words and phrases are layered on both the Tier 1 and Tier 2 choices to always account for user performance, which becomes more accurate and precise the more the user engages with the disclosed methods and systems. This layered approach ensures that the user receives the most individualized vowel-centric speech modification program possible.
[0045] As the user progresses through the vowel-centric speech modification program, the disclosed methods and systems dynamically advance the user in real-time through adaptive guidance based on the Tiers 1-3 word choices 132. In an example, the emphasis, especially over time, of the disclosed methods and systems focuses on adapting the word and oronym choices based on the user’s performance. The user performance becomes increasingly relevant to assess the user’s progress through the disclosed vowel-centric speech modification program because it is the user’s most recent speech assessment. Adapting subsequent user sessions based on the most recent user performance helps the user learn at their level and rate of content digestion rather than a fixed or structured system that is not tied to their learning progress.
[0046] FIG. 2 shows an example system 200 that helps provide vowel-centric speech modification treatment to users. The system 200 has memory 202 and multiple modules 204 that are connect to one or multiple physical processors 206. The processors 206 are electronically coupled to the modules 204 in the memory 202 to execute instructions for any one or more activated modules 204. The memory 202 can also be electronically coupled to input / output devices 208, a communications module 210, and additional memory 212 to store user data, additional modules, or other data, instructions, or information.
[0047] In the disclosed system 200 shown in FIG. 2, the modules 204 includes a speech evaluation module 214, a real-time speech analysis module 216, a visual feedback module 218, and an adaptive individualized vowel-centric speech modification program module 220. The speech evaluation module 214 evaluates a new user to the vowel-centric speech modificationAttorney Ref. No. 1033.1004. WOPCT PATENTprogram and / or re-evaluates users at various milestones throughout the program. For new users, the speech evaluation module 214 has an evaluation passage 222 that is a vowel centric reading passage, as discussed above. The real-time speech analysis module 216 evaluates user speech when the user reads the evaluation passage and when the user speaks during user sessions. For example, the real-time speech analysis module 216 receives data from a user session of the vowel-centric speech modification program and determines whether the user’s speech meets one or more evaluation criteria. Output from the real-time speech analysis module 216 is transmitted to the visual feedback module 218 that then translates the output into a visual cue for the user to adjust any determined phonological challenges. The visual feedback module 218 presents one or more visual cues to the user based on the output of the real-time speech analysis module 216 to ensure that the user has a chance to correct pronunciation by adjusting their speech according to the visual cues. The adaptive individualized vowel-centric speech modification program 220 adjust user session to the user’s specific needs.
[0048] In an example, the adaptive individualized vowel-centric speech modification program 220 uses a tiered approach to creating the individualized user sessions. As discussed above, the example can include Tiers 1-3 with Tier 1 including words and oronyms for the user session that are industry or profession specific 224. Typically, Tier 1 words and oronyms would be included in the user’s first session and / or the evaluation of the user. Tier 2 can layer on words or oronyms to the Tier 1 words and oronyms to add predictive native language transfer specific oronyms 226, as discussed above. Tier 3 layers yet another set of words or oronyms to the Tiers 1 and 2 words and oronyms that is based on adaptive user performance 228.
[0049] FIG. 3A shows an example vowel-centric speech evaluation passage 300. The vowelcentric speech evaluation passage 300 is a reading passage in Standard American English that represents all vowels equally, which means that each vowel is read the same number of times throughout the passage. The Grandfather Passage is a standardized speech evaluation passage used by SLPs to assess speech in patients. However, the conventional Grandfather Passage is consonant focused. The vowel-centric speech evaluation passage 300 shown in FIG. 3A is modified from the Grandfather Passage to be vowel-centric, which produces a vowel-centered assessment of the user’s speech. FIG. 3B shows an aviation industry specific reading passage 302 with the frequency of vowels in the passage correlating to the frequency with which the same vowels appear in aviation industry standard words and phrases, such as ICAO standardAttorney Ref. No. 1033.1004. WOPCT PATENTnumeral pronunciations, NATO phonetic alphabet terms, air traffic control commands, and high-frequency operational phrases, wherein the vowel sounds / ə / , / ɪ / , / ɛ / , / i / , / eɪ / , and / ɑ / appear at disproportionately higher frequency relative to general conversational English. In some examples, a user that pronounces a vowel with at or below a predetermined threshold accuracy, such as 50% accuracy, of the vowel's instances in either of the modified Grandfather Passage 300 shown in FIG. 3A or the aviation industry specific passage shown in FIG. 3B is considered to need treatment to help correct their phonological challenge with the vowel. The vowels requiring correction are considered a "target vowel" needing treatment. Because the aviation industry (or any other industry specific passage) focuses on weighting the vowels that are most important to the required speech in the aviation industry, the user improves the most impactful vowels first before moving on to improving less impactful vowels.
[0050] Typically, in conventional speech modification techniques, a user’s speech accuracy is evaluated based on an SLP’s trained, but subjective evaluation of the user’s speech and / or whether analyzed user input with extracted frequency values or a spectrogram produces the expected Fl and F2 frequency values. Conventionally. SLPs are not giving patients / users vowelcentric reading passages with vowels equally represented throughout the reading passage or weighted towards frequency of vowel sounds in a specific industry or setting to produce a vowelcentric speech assessment. With the modified Grandfather Passage shown in FIG. 3A or the aviation specific, weighted reading passage shown in FIG. 3B, or any vowel-centric reading passage, the user assessment is vowel-centric, which allows the vowel-centric speech modification treatment to focus on correcting vowel pronunciation on a hierarchical basis with the most valuable corrections prioritized first. Users speaking a foreign language are not considered to have a speech “disease or disorder” so the conventional consonant-centric approach is not designed to address the phonological transfer patterns that characterize accented speech. Those speaking a foreign language with an accent are not considered to have a communication disorder although they wish to improve their speech to achieve intelligibility for various reasons including professional and personal success.
[0051] Typically, a maximum of 3-4 vowels are selected to be target vowels at any given time during the user’s treatment. Treating all 18 vowels at the same time can be overwhelming for the user, which produces ineffective results. Any suitable number of vowels can be selected as target vowels. If a user has phonological challenges with all 18 vowels in Standard AmericanAttorney Ref. No. 1033.1004. WOPCT PATENTEnglish, the disclosed methods and systems select the 3-4 vowels that are most challenging for the user and most frequently occurring in speech specific to the user’s industry or profession based on a hierarchical scoring system. The scoring system can be any suitable system, such as determining whether the user’s vowel pronunciation is < 50% accurate or any other criterion.
[0052] The target vowels are also selected based on their placement in different quadrants of a vowel quadrilateral 400, as shown in FIG. 4. The vowel quadrilateral 400 helps linguists, phoneticians, and language learners visually understand how vowels differ from one another based on the physical movements of the speakers mouth opening size, tongue position, and the resulting acoustic properties (Fl and F2). The vertical axis of the vowel quadrilateral 400 corresponds to the first formant (Fl) and represents tongue height, or how open or closed the speaker's mouth is when producing a vowel. The horizontal axis corresponds to the second formant (F2) and represents the front-to-back position of the tongue in the speaker's mouth. By plotting vowels in the vowel quadrilateral 400, speakers can visually understand relationships between vowel sounds and how they might change across different languages or accents. Trained experts like SLPs with specific spectrogram interpretation training can interpret spectrograms that provide feedback on a user's vowel pronunciation, but spectrograms are difficult for an untrained person to understand. The vowel quadrilateral 400 is a visual education tool for users to help them understand how to produce a vowel sound. When a speaker produces a high vowel such as / i / as in the word "see," the tongue is positioned high in the mouth and the mouth opening is relatively closed, corresponding to a low Fl value. When a speaker produces a low vowel such as / a / as in the word "father," the tongue is positioned low in the mouth and the mouth opening is wide, corresponding to a high Fl value. When a speaker produces a front vowel such as / i / as in "see," the tongue is positioned closer to the teeth, corresponding to a high F2 value. When a speaker produces a back vowel such as / u / as in "food," the tongue is positioned farther back in the mouth, corresponding to a low F2 value. In the vowel quadrilateral 400, front vowels such as / i / and / ae / appear on the left side while back vowels such as / u / and / a / appear on the right side. High vowels such as / i / and / u / appear at the top of the vowelAttorney Ref. No. 1033.1004. WOPCT PATENTquadrilateral 400 while low vowels such as / a / and / ae / appear at the bottom of the vowel quadrilateral 400.
[0053] Vowels are difficult to perceive and produce verbally, which makes conventional approaches to accent modification focus on consonant correction that has specific placement points for the user’s tongue, lips, and mouth shape, also known as a ‘point of contact.’ Each quadrant of the vowel quadrilateral 400 represents a different tongue placement. For example, vowels in the words and oronyms given to a user for therapy can be selected from each comer of the vowel quadrilateral 400 for maximum oppositions in pronunciation. If the selected vowels for the words and oronyms given to the user are too close to each other in the vowel quadrilateral 400, the user does not process the pronunciation differences as easily. As the user progresses through a vowel-centric speech modification program, the vowels selected for the words and oronyms for the user sessions are closer in the vowel quadrilateral 400. Having a closer sound for the vowel means the user is more challenged to pronounce them in an expected manner. As the user nears the advanced and proficient levels of vowel pronunciation, the selected vowels for the words and oronyms in the user sessions may be adjacent each other in the vowel quadrilateral 400 for maximum difficulty.
[0054] FIGS. 5A and 5B are spectrograms 500A, 500B of a user that speaks two different words, respectively. All vowels have a matching, standard, and expected shape on a spectrogram that an SLP with specific training in spectrogram interpretation can locate and identify. The shape of the vowel sound in a spectrogram reflects where and how air is vibrating in a speaker’s mouth and throat. Conventionally, the specially trained SLP uses the spectrogram to create visual cues with their face and hand gestures during patient treatment to indicate to the speaker / user how to correct the pronunciation production by moving their tongue, lips, or mouth up, down, or to change the pace or pitch of the sound production. This is often accomplished by in-person or virtual treatment with demonstration by the specially trained SLP with the benefit of the spectrogram data of the changes the speaker / user needs to make. Having a user match their vocal sound production to a visual model is most effective in improving any speech production accuracy and especially to help speakers / users improve vowel pronunciation. However, because spectrograms are difficult for untrained users to interpret and even harder for them to understandAttorney Ref. No. 1033.1004. WOPCT PATENThow to alter their speech to improve their intelligibility, visual cues, supported or unsupported by a human SLP, become important to help speakers / users to improve their pronunciation.
[0055] FIG. 5A shows a spectrogram of continuous speech of a user speaking the word “he” in Standard American English. In the spectrogram 502 on the left, a native speaker of Standard American English spoke the word “he.” In the spectrogram 504 on the right, a non-native speaker of Standard American English spoke the word “he.” In both spectrograms 502, 504 and all spectrograms in general, Fl 506, 510 and F2 508, 512 refer to first and second formants, respectively, which are critical auditory components of vowels. Formants are resonant frequencies of the vocal tract that influence how different sounds are heard by others, which relates to a speaker’s intelligibility. The FIs are related to the height of the vowel sound and indicate whether the vowel sound is high or low in the speaker’s mouth. The frequency of Fl typically increases as the speaker's tongue is positioned lower in their mouth. For example, for vowel sounds like "ah" in the word "father" the frequency of the Fl is high and decreases as the tongue moves higher in the mouth to speak a vowel sound like "ee" in the word "see." The Fl reflects how "open" or "closed" the speaker's vocal tract is when they speak a given vowel sound.
[0056] The F2s are related to the “front” or “back” position of the tongue in the mouth as a speaker forms a vowel sound. Higher F2 frequencies are associated with more “front” vowels like the “I” sound in “India” while lower F2 frequencies are associated with more “back” vowels like the “oo” sound in “food.” Together, Fl and F2 provide a map of vowel sounds on the spectrogram, which helps distinguish different vowels. A low Fl 506 and a high F2 508 are typical of a front, high vowel sound like “ee” in the word “he” as it is spoken by a native speaker of Standard American English, which is depicted in the spectrogram 502 shown in FIG. 5A.
[0057] A high Fl 514 and a low F2516 are characteristic of a back, low vowel sound like “ah” in the word “audio” as it is spoken by a native speaker of Standard American English, which is depicted in the spectrogram 518 of the continuous speech of the native speaker shown in FIG.5B. The non-native speaker’s spectrogram 500A of the continuous speech as they speak the word “he” is also shown in FIG. 5A. Fl 510 (275Hz) and F2 512 (2450Hz) of the non-native speaker differ from Fl 506 (270Hz) and F2 508 (2290Hz) of the native speaker by an increased distance between the first and second formant frequencies, wherein Fl 510 (275Hz) of the non-native speaker is comparable in frequency to Fl 506 (270Hz) of the native speaker and F2 512Attorney Ref. No. 1033.1004. WOPCT PATENT(2450Hz) of the non-native speaker is higher in frequency than F2 508 (2290Hz) of the native speaker, reflecting a more tense and fronted vowel production in which the non-native speaker's tongue is positioned further forward than the target, producing a vowel that overshoots the expected acoustic range for the target vowel in Standard American English.
[0058] Similarly, the non-native speaker’s spectrogram 520 as they speak the word “audio” is shown in FIG. 5B. Fl 522 (730 Hz) and F2524 (1350 Hz) of the non-native speaker differ from Fl 514 (710 Hz) and F2516 (1090 Hz) of the native speaker by a deviation in formant frequency values, wherein Fl 522 (730 Hz) of the non-native speaker is higher in frequency than Fl 514 (710 Hz) of the native speaker and F2 524 (1350 Hz) of the non-native speaker is higher in frequency than F2 516 (1090 Hz) of the native speaker, indicating a centralized vowel production in which the non-native speaker's vowel is produced with a more forward tongue position than the native speaker's target.
[0059] The Fl 506, 510 in the spoken word “he” and the Fl 514, 522 in the spoken word “audio” respectively give the disclosed methods and systems information about the vertical or height dimension of the speaker’s mouth opening and tongue height utilizing formant frequencies. The F2 508, 512 in the spoken word “he” and the F2 516, 524 in the spoken word “audio,” respectively give the disclosed methods and systems information about the horizontal or front to back dimension of the speaker’s tongue placement. The Fl and F2 frequencies and their shapes usually refers to formant transitions (how the formant frequency changes over time during the vowel) of the non-native speaker and are used to create a visual cue that the speaker must match to improve their intelligibility. This technique results in the non-native speaker’s FIs and F2 appearing more like the expected FIs and F2s frequencies and shape of the native speaker. Because a native speaker is oftentimes unavailable to model expected vowel sounds with visual cues for a non-native speaker, the vertical and horizontal dimensions of the non-native speaker’s FIs and F2s on their spectrograms are translated into a visual output with one or more visual cues that mimics the visual output or feedback that would be provided by a native speaker or SLP training a non-native speaker using their face and hand gestures.
[0060] As shown in FIG. 5C, the vowel-centric speech modification system determines Fl and F2 frequency values without spectrograms 500C although spectrograms can optionally be used as a supplement to validate data or separately by a human SLP specially trained in spectrogramAttorney Ref. No. 1033.1004. WOPCT PATENTinterpretation. When the user inputs data to the system, such as through a reading passage or speech sample or in real-time as they progress through a user session of guided speech,.the vowel-centric speech modification system identifies vowel sound segments and their acoustic properties in the user data and translates it into correlating Fl and F2 values by applying formant estimation to the acoustic signal to extract the resonant frequency values corresponding to each identified vowel sound. These values can be tracked in real-time as the user speaks to create a continuous stream of Fl and F2 values without requiring the system to create spectrograms. This continuous stream of Fl and F2 values provides the boundaries of the visual output the user then sees on their device. That visual output is similar to the visual cues an SEP specially trained in spectrogram interpretation would provide the user after analyzing their spectrogram without the need for either the spectrogram itself or the SEP visual cues.
[0061] Correlating the Fl and F2 values with the output visual output gives the user the benefit of understanding the visual cues that improve learning capacity and efficiency while remaining in a range of error free learning and guided corrective learning of the vowel sounds. Without the visual cues produced by the Fl and F2 frequencies that are extracted from the user input of their speech, the user could be practicing sounds without the benefit of the powerful visual cues that bolster learning comprehension and increase the speed with which accent modification occurs to and beyond a proficient level. In this way, FIG. 5C shows a method of creating visual output based on user input in real-time 500C without using a spectrogram. The user speaks a reading passage, provides a speech sample, or speaks in real-time, such as during guided speech of a user session 526. The vowel-centric speech modification program extracts the Fl and F2 frequency values from the user input of speech 528 by applying formant extraction to the acoustic signal to extract the resonant frequency values corresponding to each identified and targeted vowel sound. The vowel-centric speech modification program then correlates the extracted Fl and F2 frequency values with visual output 530, such as one or more visual cues like an animation, object in motion, alerts, visual indicator(s), and the like. The method adjusts the visual output in real-time in response to the changing real-time user input of speech 532 to provide the user with a dynamically adjusting visual output that tracks their speech production in real-time so they can make real-time corrections with the benefit of the visual output as feedback.
[0062] FIG. 6 shows a chart of the expected frequencies of the vowel sounds in Standard American English 600. This chart can be used as the standard or threshold values for comparingAttorney Ref. No. 1033.1004. WOPCT PATENTthe Fl and F2 frequency values of the user’s input data of speech. When the extracted Fl and F2 frequency values of their speech does not align with this chart, for example, then the visual output indicates an error with various visual cues, such as changing an object the user is tracking (e.g., an aircraft or flight path in an aviation example), outputting an alert or arrows that indicate the necessary change in the user’s pronunciation, or the like. Column 602 shows the vowel sound with an example word for clarity. Column 604 shows the International Phonetic Alphabet (IPA) symbol associated with the vowel sounds shown in column 602. Column 606 shows the International Civil Aviation Organization (ICAO) word for a particular aviation industry related word or action. The ICAO is a United Nations organization that manages the administration and governance of aviation safety, security, efficiency, and environmental protection across 193 member countries across the globe. Column 608 shows the Federal Aviation Administration (FAA) word for the equivalent of the ICAO word shown in Column 606. The FAA is the United States correlate aviation safety and regulatory agency to the ICAO. Column 610 includes the Air Traffic Control (ATC) word for the equivalent of the ICAO and FAA words shown in Columns 606, 608, respectively. Column 612 lists the average Fl frequency in Hz for each respective vowel sound. Column 614 lists the average F2 frequency in Hz for each respective vowel sound.
[0063] The system determines whether a speaker's produced vowel formant values fall within a target corridor defined by a statistically derived range encompassing approximately 80% of native English speakers. Specifically, for each target vowel, the system extracts the first formant frequency (Fl) and second formant frequency (F2) from the speaker's speech signal and compares these values against a predetermined acceptable range. The acceptable range for each formant is computed as the mean value ± 1.28 standard deviations from a reference population of native English speakers, thereby capturing approximately 80% of the native speaker distribution. A speaker's production of a given vowel is classified as intelligible when both the Fl and F2 values fall within their respective acceptable ranges simultaneously. This chart can be used to compile a list of industry specific words for the aviation industry in both the US and globally and help determine the expected vowel sounds along with expected frequencies of the Fl and F2 that are used to evaluate the intelligibility of the pilots applying for licensure. In alternative embodiments, the threshold for evaluating the formant values may differ, such as requiring the vowel formant values to fall within a higher or allowing them to be in a lower percentage of correct speech by native English speakers or any other foreign language the user is learning. TheAttorney Ref. No. 1033.1004. WOPCT PATENTstandard deviation from a reference point could also be larger or smaller, depending on the system needs. For example, the native speaker distribution could be within a range of 90% accuracy of a native English speaker while the standard deviation is smaller for those users learning the language at a higher level, such as those that want to achieve conversational status with little to no accentedness.
[0064] FIGS. 7A and 7B show aviation industry specific vocabulary categories 700A, 700B that help the disclosed methods and systems select Tier 1 words and oronyms for industry or profession specific choices. These aviation specific categories can be included in any one or more of the ICAO, FAA, and ATC specific words described above and shown in FIG. 6. In the aviation industry, for example, the following categories of terms are essential to pilots performing their jobs: ATC essentials; numbers and altitudes; emergency terms; standard phrases; clearances and routing; VFR and flight status; restrictions; transponder and radio; surface movement and ground; pattern and VFR tower; aircraft approach and landing; weather and runway conditions; emergency priority; dispatch and planning; performance and weights; flight deck standard operating procedures; datalink and oceanic; low-visibility and weather operations; ramp and surface operations; maintenance and airworthiness; and cabin and crew coordination. Other industries or professions and different social situations have different categories of words or oronyms and different words entirely that are specific to their use case. Each specific word or phrase in the industry / profession / social environment specific environment has a phonetic pronunciation that enables comprehensive El specific target language training for the speakers, which can be integrated into the words and oronyms practiced by the users. In some examples, the users can practice only industry or profession specific words and oronyms while in other examples, they can practice a mix of industry or profession specific words and oronyms and other conversational words and oronyms.
[0065] FIGS. 7A and 7B show words that correlate to the aviation industry. However, oronyms can also be added or substituted when they are helpful. Oronyms are a word or phrase in the user’s El that correlates to the words or phrase needing correction in the foreign language they are trying to learn. This provides a forced prosodic alignment between the two languages so the user is understood in the foreign language while relying on familiar pronunciation skills and speech sounds they already have in their LE Such forced prosodic alignment identifies sounds the user already has from their LI and identifies them in the target words or phrases even if theyAttorney Ref. No. 1033.1004. WOPCT PATENThave different meanings. For example, a Spanish speaker may be learning how to speak the number seven “7” in English by looking at the oronym “se ven” in Spanish, which means “they see each other” in a direct English translation. This oronym has a different definition in English compared to Spanish, but the pronunciation of “se ven” in Spanish is intelligible in English as the number 7. This quickens the learning process for the user by relying on sounds and phrases they know well instead of starting the learning process from the basic level. These familiar correlations are included in the selected target words, phrases, and the like based on the user’s LI, their reading passage, their speech samples, and their real-time speech because oronyms correlations induces correct foreign language (L2) prosodic output through LI phonological transfer.
[0066] In the disclosed methods and systems, the visual feedback the speaker or user sees is embodied in an application or other computer-readable instructions on their mobile device or computer. In an example, the disclosed methods and systems are embodied in a mobile device application 800, 900, as shown in FIGS. 8 and 9. FIG. 8 shows visual feedback in the form of a graphical user interface (GUI) output with a dynamically changing graphic of an aircraft 804 that the speaker must maintain within a flight path of intelligibility 806. Alternatively, the aircraft 804 could remain stationary while the flight path changes in shape, contour, or size by minimizing or expanding it. Either the aircraft 804 or the flight path changes while the other remains stationary so there is a point of reference for the changes the user needs to make relative to a specific baseline standard of proficiency or higher level learning, as desired.
[0067] In this example, the flight path of intelligibility 806 is a tubular shape with exterior surface boundaries that correlate to the Fl and F2 frequencies in a spectrogram of expected frequencies of a native speaker of various words given to a non-native speaker or user of the application to practice. The lower exterior surface 808 of the tubular flight path of intelligibility 806 correlates to the Fl frequency of a spectrogram of live speech spoken by the user practicing in the user session. The upper exterior surface of the tubular flight path of intelligibility 810 correlates to the F2 frequency of a spectrogram of live speech spoken by the user practicing in the user session. These outer boundaries 808, 810 can be the outer limits of intelligibility for proficient pronunciation of the vowel sounds. Inner boundaries 812, 814 that create a tubular flight path of intelligibility with a smaller diameter can represent a higher skill level and accuracy of the vowel pronunciation by the user. For example, the second boundary 812Attorney Ref. No. 1033.1004. WOPCT PATENTcorrelates to an Fl and F2 frequency that indicates the speaker is highly skilled in vowel pronunciation while the third innermost boundary with the smallest diameter correlates to an Fl and F2 frequency that indicates the speaker has mastered conversational pronunciation of the vowels.
[0068] The flight path of intelligibility 806 has a shape or pathway that corresponds to the expected up / down and front / back movements of the speakers tongue required for the user’s speech to remain intelligible over various thresholds (ex: proficient; highly skilled; conversational mastery). This pathway is modeled expected Fl and F2 values of the words spoken by a native speaker. When the user speaks the selected words during a user session, the Fl and F2 values on the spectrogram are translated in real-time into a user flight path 816 for the aircraft 804 shown on the GUI. The shape and contour of the user flight path 816 follows the real-time Fl and F2 frequencies of the user’s continuous speech. When the user’s speech is within the expected Fl and F2 frequency values of proficient vowel pronunciation, the user’s flight path appears are a green and / or with check marks 818 or other visual cues visually indicating to the user that their pronunciation is within an expected range. When the user’s speech is outside the expected Fl and F2 frequency values of proficient speech, the user’s flight path appears as red or with an alarm or alert visual cue 820 to visually indicate that the user is outside the expected Fl and F2 frequency values of proficient speech. When the user crosses the threshold from the expected Fl and F2 frequency values to unexpected Fl and F2 frequency values an alert 822 can appear with instructions on how to correct the user’s pronunciation back to the expected Fl and F2 frequency values.
[0069] For example, FIG. 8 outputs an instructive alert 822 on the GUI when the user’s speech is outside the expected Fl and F2 frequency values 808, 810. The instructive alert 822 includes one or more indicators of what the user needs to do to correct their pronunciation. For example, the instructive alert 822 includes a “stall warning” to point the nose of the aircraft 804 down, which would cause the aircraft 804 to return to the expected Fl and F2 frequency values 808, 810 of the flight paths of intelligibility 806. The instructive alert 822 shown in FIG. 8 also includes specific instructions on how to accomplish the correction by having the user lower the pitch pronunciation of the vowel sound. In the example shown in FIG. 8, the combination of the position of the aircraft 804 with its nose outside the flight path of intelligibility 806; the instructive alert with instructions on how to change the user’s pronunciation; and the arrows 826Attorney Ref. No. 1033.1004. WOPCT PATENTindicating to lower the pitch pronunciation of the user’s vowel sounds provide the user with visual output on how to correct their pronunciation to keep the aircraft 804 within the flight path of intelligibility 806.
[0070] The disclosed methods and systems determine the Fl and F2 frequency values 808, 810 based on the user speaking a passage, such as the sample text 824 shown in FIG. 8, and creating a spectrogram of the user’s speech. The disclosed methods and systems identify the Fl and F2 frequencies on the user’s spectrogram based on known Fl and F2 frequency patterns of native and non-native speakers. Those identified Fl and F2 frequency values from the spectrogram then become the shape and contour of the user’s flight path by constricting or expanding based on the user’s input. In some examples, such as the output shown in FIG. 8, the methods and systems also output another visual cue, such as the arrows 826 shown in FIG. 8 to indicate the corrective action the user needs to take. In the example shown in FIG. 8, the arrows 826 indicate that the user needs to lower their F2 frequency value.
[0071] The aircraft moving along a flight path is a visual tool for users to follow when speaking and making speech corrections. This example is specific to pilots in the aviation industry. However, similar visual feedback specific to any other industry or profession or social situation of interest can be created for another use case.
[0072] FIG. 9 shows another embodiment of visual feedback in the form of a graphical user interface (GUI) output with a dynamically changing graphic of an aircraft 902 that the speaker must maintain within a flight path of intelligibility 904. In this example, the flight path of intelligibility 904 is circular. The user must maintain the aircraft 902 within the circle. The lower exterior boundary 906 of the flight path of intelligibility circle correlates to the Fl frequency value of the expected vowel sounds, and the upper exterior boundary 908 of the circle correlates to the F2 frequency value of the expected vowel sounds. Like the example shown in FIG. 8, the methods and systems shown in FIG. 9 have multiple levels of intelligibility of the user’s speech. The outer ring made of the lower exterior boundary 906 and the upper exterior boundary 908 are the minimum threshold that indicates a proficient level of intelligible speech spoken by the user. The middle ring 910 indicates a high level of speech by the user while the innermost ring 912 indicates conversational mastery of the vowel sounds by the user. When the user’s speech is outside the expected range of Fl and F2 values, the aircraft 902 veers outside theAttorney Ref. No. 1033.1004. WOPCT PATENTrings 906, 908; 910; 912, which triggers visual output, such as the red quadrant 914 illuminating and the instructive alert 916 that visual outputs text indicating the area of correction of the user’s speech and an instruction on how to correct it. Other suitable visual output can be added or some of the described visual output can be removed.
[0073] The foregoing description, for purposes of explanation, used specific nomenclature to provide a thorough understanding of the disclosure. However, it will be apparent to one skilled in the art that the specific details are not required in order to practice the systems and methods described herein. The foregoing descriptions of specific embodiments or examples are presented by way of examples for purposes of illustration and description. They are not intended to be exhaustive of or to limit this disclosure to the precise forms described. Many modifications and variations are possible in view of the above teachings. The embodiments or examples are illustrated and described in order to best explain the principles of this disclosure and practical applications, to thereby enable others skilled in the art to best utilize this disclosure and various embodiments or examples with various modifications as are suited to the particular use contemplated. It is intended that the scope of this disclosure be defined by the following claims and their equivalents.
Claims
Attorney Ref. No. 1033.1004. WOPCT PATENTWhat is claimed is:
1. A method of modifying speech using a vowel-centric speech modification technique in a target language, comprising:receiving input of a vowel-centric reading passage sample for a user;determining a first formant (Fl) and a second formant (F2) in the input of the vowelcentric reading passage sample, both of the Fl and F2 measured in frequency values;identifying target intelligibility vowels in which one or both of the Fl frequency values or F2 frequency values of the input of the vowel-centric reading passage sample do not meet a respective Fl intelligibility criterion or an F2 intelligibility criterion;individually selecting target vowels, vowel combinations, words, phrases, sentences, or conversational discourse topics from the identified target intelligibility vowels, each selected target word related to speech intelligibility in the target language;creating an individualized user vowel-centric speech modification program based on the individually selected target vowels, the individualized user vowel-centric speech modification program including multiple user sessions;creating a visual output with a first visual cue related to the Fl and a second visual cue related to the F2;receiving input from a first of the multiple user sessions, the input dynamically changing the first visual output related to the Fl and the second visual output related to the F2, respectively, based on Fl frequency values and F2 frequency values of the first of the multiple user sessions, respectively.
2. The method of claim 1, wherein the vowel-centric reading passage equally represents or weights each vowel in the target language based on the respective vowel frequency occurrence in speech associated with an industry or profession of the user.
3. The method of claim 1, wherein the target language is Standard American English.
4. The method of claim 1, wherein the vowel-centric reading passage sample is modified to focus on vowel intelligibility for an aviation specific passage or industry specific passage.Attorney Ref. No. 1033.1004. WOPCT PATENT5. The method of claim 1, wherein the categorized group of target vowels, vowel combinations, words, phrases, sentences, or conversational discourse topics relate to an industry or a profession.
6. The method of claim 5, wherein the industry or profession is aviation.
7. The method of claim 1, wherein creating the individualized user vowel-centric speech modification program includes individually selected and created oronyms having the individually selected target vowels.
8. The method of claim 7, wherein the individually selected and created oronyms are individually selected safety critical oronyms related to an industry or profession.
9. The method of claim 1, wherein creating the individual user vowel-centric speech modification program is based on an industry or a profession of the user.
10. The method of claim 9, wherein creating the individual user vowel-centric speech modification program is also based on typical transfer patterns of a native language of the user.
11. The method of claim 10, wherein creating the individual user vowel-centric speech modification program is further based on user performance included in input from one or more of the multiple user sessions.
12. The method of claim 11, wherein input from user performance of a first of the multiple user sessions is compared to a threshold or set of criteria, and the user vowel-centric speech modification program is dynamically changed in a subsequent of the multiple user sessions based on the input of the user performance of the first of the multiple user sessions.
13. The method of claim 1, further comprising creating the input from user performance in each session of the multiple user sessions of the user vowel-centric speech modification program to be progressively more difficult, and further comprising determining that the input from the user performance meets or exceeds a threshold or set of criteria for one of theAttorney Ref. No. 1033.1004. WOPCT PATENTmultiple user sessions before allowing the user to progress to a subsequent of the multiple user sessions.
14. The method of claim 1, further comprising creating input from user performance in each session of the multiple user sessions of the user vowel-centric speech modification program to be progressively more difficult, and further comprising requiring the user to repeat one of the user sessions if the input from the user performance does not meet or exceed a threshold or set of criteria from the one of the user sessions.
15. The method of claim 1, wherein the visual output relates to a profession of the user, and the first visual cue and the second visual cue relate to an aspect of the profession of the user.
16. The method of claim 15, wherein the profession is aviation and the first visual cue and the second visual cue relate to maintaining an airplane within a flight path with boundaries that correlate to F1 and F2, respectively.
17. The method of claim 1, wherein the creating the visual output related to the first visual cue and the second visual cue is created in real-time.
18. A system of modifying speech using a vowel-centric speech modification technique in a target language, comprising:an input configured to receive input of a vowel-centric reading passage sample for a user; a processor configured to:determine a first formant (F1) and a second formant (F2) in the input of the vowel-centric reading passage sample, both of the F1 and F2 measured in frequency values;identify target intelligibility vowels in which one or both of the F1 frequency values or the F2 frequency values of the spectrogram of the vowel-centric reading passage sample do not meet a respective F1 intelligibility criterion or an F2 intelligibility criterion;individually select target vowels from the identified target intelligibility vowels, each selected target vowel related to speech intelligibility in the target language;create an individualized user vowel-centric speech modification program based on the individually selected target vowels, the individualized user vowel-centric speech modification program including multiple user sessions; andAttorney Ref. No. 1033.1004. WOPCT PATENTan output configured to output a first visual cue related to the F1 intelligibility criterion and a second visual cue related to the F2 intelligibility criterion based on user performance during one of the multiple sessions of the individualized user vowel-centric speech modification program.
19. The system of claim 18, wherein the processor is further configured to create the individualized user vowel-centric speech modification program to include individually selected oronyms having the individually selected target vowels.
20. The method of claim 19, wherein the individually selected oronyms are individually selected safety critical oronyms related to an industry or profession.
21. The method of claim 20, wherein the processor is further configured to create the individualized user vowel-centric speech modification program based on typical transfer patterns of a native language of the user.
22. The method of claim 21, wherein the processor is further configured to create the individualized user vowel-centric speech modification program based on user performance included in input from one or more of the multiple user sessions.
23. The method of claim 18, wherein the visual output relates to a profession of the user, and the first visual cue and the second visual cue relate to an aspect of the profession of the user.
24. The method of claim 18, wherein the output is configured to output the first visual cue and the second visual cue in real-time based on the user performance during one of the multiple sessions of the individualized user vowel-centric speech modification program.
25. A method of modifying speech using a vowel-centric speech modification technique in a target language, comprising:receiving real-time user speech as input for a user;determining a first formant (F1) and a second formant (F2) in the real-time user speech input, both of the F1 and F2 measured in frequency values;Attorney Ref. No. 1033.1004. WOPCT PATENTidentifying target intelligibility vowels in which one or both of F1 frequency values or F2 frequency values of the real-time user speech do not meet a respective F1 intelligibility criterion or F2 intelligibility criterion;creating a visual output in response to the real-time user speech with a first visual cue based on the Fl frequency values and a second visual cue based on the F2 frequency values; anddynamically and continuously changing the first visual output related to the F1 and the second visual output related to the F2, respectively, based on changing F1 frequency values and F2 frequency values of the real-time user speech.