Voice audiometry score
By providing test audio of target words through computing devices and performing phoneme analysis, the problem of difficulty in assessing users' hearing ability in existing audiometry technologies is solved, enabling personalized adjustment and rehabilitation training of hearing devices and improving the therapeutic effect of hearing devices.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-11-18
- Publication Date
- 2026-03-03
AI Technical Summary
Existing audiometry technologies are insufficient to effectively assess a user's hearing ability, especially when enhanced by assistive hearing devices, and cannot accurately identify a user's hearing problems to provide personalized treatment.
The device provides test audio of the target word, receives the user's response, converts it into a speech representation, uses phoneme analysis and comparison algorithms to assess the user's hearing ability, generates a phoneme score, and then determines the adjustment and rehabilitation training plan for hearing devices.
It enables accurate assessment of users' hearing abilities, provides personalized hearing device adjustments and rehabilitation training, and improves the therapeutic effect of hearing devices.
Smart Images

Figure CN114727780B_ABST
Abstract
Description
[0001] This application was filed as a PCT international patent application on November 18, 2020, and claims priority and benefit to U.S. Provisional Application Serial No. 62 / 938,572, filed on November 21, 2019, the entire disclosure of which is incorporated herein by reference in its entirety. Background Technology
[0002] Audiometry involves measuring a person's hearing. Audiometry tests may include measuring a user's response to sound. For example, a user may listen to short words in a soundproof room and verbally repeat what they understand. In unassisted audiometry, the user's natural hearing (e.g., without auditory enhancement) is tested. In assisted audiometry, the user's hearing is tested with auditory enhancement devices. Tests may be performed on one or both ears of the user at a time. The content of the test may include monosyllabic words, disyllabic words, entire sentences, or other forms. Audiologists may listen to the responses and score them. Summary of the Invention
[0003] In the example, there is a computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: select a target word from a plurality of words in an audiometry test; provide test audio of the pronounced target word; obtain a user response as text data via a user interface; convert the text data into a speech representation of the user response; obtain a speech representation of the target word; compare the user response speech representation and the speech representation of the target word to obtain comparison data; and determine the user's hearing ability based on the comparison data.
[0004] In another example, there is a method that includes: providing a user with test audio including a target word pronounced; obtaining a user response from the user; converting the user response into a user response speech representation; converting the target word into a target word speech representation; comparing the user response speech representation and the target word speech representation to obtain comparison data; and determining the user's hearing ability based on the comparison data.
[0005] In yet another example, there exists a system comprising: a text input unit; a sound output unit; one or more processors; and a memory storing instructions that, when executed by the one or more processors, cause the one or more processors to: provide test audio including a pronounced target word through the sound output unit; obtain text data from the text input unit; convert the text data into a user-response speech representation; obtain a target word speech representation of the target word; compare the user-response speech representation and the target word speech representation to obtain comparison data; and determine the user's hearing ability based on the comparison data. Attached Figure Description
[0006] In all the accompanying drawings, the same reference numerals denote the same elements or elements of the same type.
[0007] Figure 1 An audiometry system for performing one or more audiometry tests is shown.
[0008] Figure 2 The method for conducting audiometry tests is shown.
[0009] Figure 3 An example of the final score generated based on the response to the target word is shown.
[0010] Figure 4 Another example of a final score generated based on the response to the target word is shown.
[0011] Figure 5 is from Figure 5A , Figure 5B and Figure 5C The composition shows the results of an example audiometry test.
[0012] Figure 5A Example audiometry test results are shown.
[0013] Figure 5B An example implementation of the results table is shown.
[0014] Figure 5C An example implementation of phoneme analysis data is shown.
[0015] Figure 6 An example cochlear implantation system is shown that can benefit from the use of the techniques disclosed herein.
[0016] Figure 7 This is a view of an example percutaneous bone conduction device from which one could benefit from the use of the techniques disclosed herein.
[0017] Figure 8 An example of a computing system that can be used to implement one or more of the disclosed examples is shown. Detailed Implementation
[0018] This application relates to audiometry testing techniques. For each target word in an audiometry test of one or more target words, the testing system can audibly provide the target word to a user and receive a response to the content perceived audibly by the user. The response may include glyphs (e.g., the orthographic spelling of the content perceived by the user). The glyphs are then converted into phonemes using, for example, a language dictionary or pronunciation rules to produce one or more phoneme representations of the response. A comparison can then be made between the target word and the speech representation of the user response (e.g., using the Levinshtein algorithm to determine the distance between them). In the presence of multiple potential speech representations, multiple comparisons can be made (e.g., comparing each user response representation with each target word representation), and the best matching pair is used.
[0019] The results of one or more comparisons can be used to determine a phoneme score. Other data can be used to supplement or replace the phoneme score. Error patterns can be identified across multiple comparisons in an audiometry test. For example, someone may be prone to making specific types of errors, such as difficulty hearing vowels, fricatives, or plosives. The resulting phoneme score can reflect a user's hearing ability and, if the results indicate that the user may be suffering from hearing loss, can serve as the basis for one or more treatment measures. For example, treatment measures may include providing the user with hearing aids or modifying the user's existing hearing aids (e.g., adjusting the low-frequency gain of the device in response to vowel errors or adjusting the high-frequency gain in response to fricative errors). Other treatment measures may include using rehabilitation exercises, such as playlists containing words that the user cannot hear clearly.
[0020] As a specific example, the target word could be "dog," and the user's response could be "do." The target word is then converted into three different pronunciations: Each pronunciation is related to a different pronunciation of the word. The user's response ("do") is then phonetically represented (e.g., using a word-to-phoneme system) into three possible speech forms: And / do: / . The resulting phonemes can be compared, with the best match falling between the target word's phonetic representation and the target word's phonetic representation. User response voice representation Between. Comparing these two representations yields a phoneme score of 67% (e.g., two out of three parts match).
[0021] Figure 1 An example audiometry system for performing audiometry tests is shown in the figure.
[0022] Audiometry system
[0023] Figure 1An audiometry system 100 is illustrated for performing one or more audiometry tests on a user. The audiometry system 100 includes a computing device 110 configured to perform audiometry. The audiometry system 100 may also include a server 170 connected to the computing device 110 via a network 102. The audiometry system 100 may also include a clinician computing device 180. In some examples, the user of the audiometry system 100 may be a recipient of a hearing device 120.
[0024] Network 102 is a computer network, such as the Internet, which facilitates electronic communication of data between computing devices connected to network 102.
[0025] Computing device 110 may be a device with computer functionality. Computing device 110 may be a consumer computing device owned or primarily used by a user or their parent or caregiver, such as a telephone, tablet, laptop, desktop computer, consumer augmented reality device, consumer virtual reality device, smartwatch, or consumer heart rate monitor. In other examples, computing device 110 may be a device owned or used by an organization such as a school, clinic, or hearing service provider. As shown, computing device 110 may include a text input component 112, a sound output component 114, and a hearing measurement application 116. If the user has a hearing device 120, computing device 110 may include a hearing device application 118. Computing device 110 may include devices related to hearing devices. Figure 8 The computing system 800 describes one or more components or functions.
[0026] Text input component 112 may be a component of computing device 110 through which it receives text data from a user, such as a touchscreen configured to receive touch input (e.g., computing device 110 can convert touch input into text data). In the case where text input component 112 includes a touchscreen, text input component 112 may display a virtual keyboard and convert tap or swipe touch input at virtual keyboard keys into text data. In another example, text input component 112 is one or more buttons (e.g., keyboard keys) configured to receive button input and convert button input into text data. In yet another example, text input component 112 is a component of a speech-to-text system provided by computing device 110 (e.g., a microphone). Computing device 110 may be configured such that applications running on computing device 110 can receive text input via text input component 112.
[0027] The audio output component 114 may be a component through which the computing device 110 provides audio output, such as one or more speakers. The audio output component 114 may be a component that collaborates with another device to provide audio output. For example, the audio output component 114 may be an audio output port of the computing device 110 or a wireless transmitter (e.g., a Bluetooth component) configured to connect to another device (e.g., a speaker, headphones, or a component of the hearing device 120) that provides audio output. The computing device 110 may be configured such that applications running on the computing device 110 can provide audio output via the audio output component 114.
[0028] The audiometry application 116 may be software that operates on computing device 110 and causes the execution of the audiometry operations described herein. In many examples, the audiometry application 116 provides audio output, receives user responses, analyzes the responses, and provides an indication of the user's hearing ability. Figure 2 Example operations performed by the audiometry application are described herein. In some examples, the instructions may be available as part of a downloadable package, such as one available from a software distribution platform. In some examples, the audiometry application 116 is accessed via a browser that provides audiometry functionality to the access server 170 (e.g., an audiometry test web application may be provided by the server 170 and accessed by a browser on the computing device 110).
[0029] The hearing device application 118 may be software that operates on the computing device 110 and collaborates with the hearing device 120. The hearing device application 118 may be stored as computer-executable instructions in the memory of the computing device 110, which, when executed, perform one or more tasks related to the hearing device application 118. For example, the hearing device application 118 may control the hearing device 120 (e.g., by automatically or based on input received from a receiver at the computing device 110 to modify hearing device settings 122), monitor the use of the hearing device 120, and obtain data from the hearing device 120. The computing device 110 may be connected to the hearing device 120 via, for example, a wireless radio frequency communication protocol (e.g., Bluetooth or Wi-Fi). The hearing device application 118 may transmit or receive data from the hearing device 120 through such a connection. The hearing device application 118 can be used to stream audio, for example, from the microphone of the computing device 110 or from an application running on the computing device 110 (e.g., the hearing measurement application 116), to the hearing device 120.
[0030] The hearing device 120 can be a device related to a user's hearing. The hearing device 120 can take various forms, including cochlear implants, electroacoustic devices, percutaneous bone conduction devices, passive percutaneous bone conduction devices, active percutaneous bone conduction devices, middle ear devices, fully implantable hearing devices, tinnitus management devices, mostly implantable hearing devices, auditory brainstem implants, hearing aids, dental-anchored hearing devices, personal sound amplification products, other auditory prostheses, and combinations of the above devices (e.g., a binaural system comprising a prosthesis for the recipient's first ear and the same or a different type of prosthesis for the second ear). Specific example embodiments of the hearing device 120 are described in [details omitted]. Figure 6 (showing a cochlear implant) and Figure 7 (A percutaneous bone conduction device is shown) is described in more detail. The techniques disclosed herein can be used with sensory devices such as consumer hearing devices (e.g., hearing aids or personal sound amplification products). Furthermore, individuals with normal hearing or hearing loss can use hearing device 120, such as headphones or a standard audiometer, during hearing tests. As shown, hearing device 120 may include a sound output component 114. Hearing device 120 can operate according to one or more hearing device settings 122.
[0031] The auditory device settings 122 may be one or more parameters having values that influence how the auditory device 120 operates. For example, the auditory device 120 may receive audio input from the environment (e.g., using a microphone), convert the audio input into a stimulus signal, and use the stimulus signal to generate stimulation (e.g., vibration or electrical stimulation) to elicit auditory perception in a user. The auditory device settings 122 may include a mapping having minimum and maximum stimulus levels for the stimulation channels. The auditory device 120 can then use this mapping to control the amount of stimulation provided. In the case that the auditory device 120 is a cochlear implant, the mapping may influence which electrodes of the cochlear implant are used for stimulation and how much based on the received audio input. In some examples, the auditory device settings 122 include two or more predefined groups of settings that can be selected by the recipient. The auditory device settings 122 may also include settings that modify the sensory input before it is converted into a stimulus signal. Such settings may include, for example, specific audio equalizer settings that can increase or decrease the audio intensity at various frequencies. In the example, auditory device settings 122 may include a minimum threshold for the received audio input to elicit stimulation, a maximum threshold for preventing stimulation above a level that would cause discomfort, gain parameters, intensity parameters (e.g., loudness), and compression parameters. Auditory device settings 122 may include settings that affect the dynamic range of the stimulation produced by auditory device 120. As described above, many auditory device settings 122 affect the physical operation of auditory device 120, such as how auditory device 120 provides stimulation to a user in response to audio input received from the environment. Therefore, modifying auditory device settings 122 can modify the therapy provided by auditory device 120. Examples of settings, setting modifications, and preprocessing for auditory prostheses are described in U.S. Patent Nos. 9,473,852 and 9,338,567, which are incorporated herein by reference for any and all purposes.
[0032] Server 170 may be a server computing device located remotely from computing device 110. Server 170 may include processing units and memory, which... Figure 8The following describes the process in more detail. Server 170 may also include instructions executable to perform one or more of the operations described herein. Server 170 may be communicatively coupled to computing device 110 via network 102. Server 170 may be indirectly communicatively coupled to hearing device 120 via computing device 110 (e.g., via hearing device application 118). In some examples, computing device 110 may be considered a client device of server 170. In some examples, functionality provided by server 170 or its components may be provided by or located on a device local to the recipient (e.g., computing device 110 or hearing device 120). One or both of hearing device application 118 and audiometry application 116 may be client applications configured to interact with server 170. For example, as shown, server 170 may include audiometry data 172, such as audiometry test data that can be used to provide audiometry tests. The test data may specify the types of tests to be run and the data used for those tests. For example, the test data may include audio files of spoken target words. The audio file can then be streamed to computing device 110 or provided to computing device 110 as download content. In some examples, audiometry data 172 includes the results of audiometry tests, for example, for review by a clinician via a clinician's computing device. In examples, audiometry test processing can be performed at a server. For example, one or more operations described herein as being performed on computing device 110 by audiometry application 116 can be performed at least partially on server 170.
[0033] The clinician computing device 180 can be a computing device used by a clinician. A clinician can be a medical professional, such as an audiologist. In this example, the clinician is a medical professional who provides care or supervision to a user. The clinician computing device 180 includes one or more software programs that can be used to monitor audiometry tests. For example, in response to an audiometry test performed for a user, the test results can be provided to the clinician via the clinician computing device 180 for analysis or to determine how to proceed. In some examples, the clinician can use the clinician computing device 180 to remotely connect to the computing device 110 to perform audiometry tests.
[0034] method
[0035] Figure 2 An example method 200 for performing an audiometry test 202 is shown. In some examples, method 200 may be performed as a result of one or more instructions being executed by one or more processors. For example, the instructions may be audiometry application instructions stored on a non-transitory computer-readable medium such as one of the computing devices 110.
[0036] The hearing test 202 can be a test of a user's hearing ability. The hearing test 202 can take any of a variety of forms. The hearing test 202 can include a list of target words 204 to be audibly presented to the user. Each target word 204 can be a specific cue to be presented to the user and to which the user provides a response (in many examples, an instruction to the user on what the target word should be understood). For example, the target word 204 can be a consonant-vowel-consonant (CVC) word, a CCVC word, a CVCC word, or another form. A CVC word is not necessarily just a three-letter word. Instead, it can be a word with a consonant, followed by a vowel, and then another consonant. Although mentioned in the context of "word" in this document, in some examples, the target word 204 can be a phrase (with multiple words) or part of a word.
[0037] The audiometry test 202 can specify a list of target words 204 in a suitable format, such as a markup language (e.g., XML). In some implementations, a predefined list of target words 204 is not required. In some examples, the contents of the list of target words 204 are arbitrarily selected. In other examples, the list or its contents may be selected for a specific reason, such as testing a particular type of word or sound that is difficult or considered difficult for the user to perceive. In yet another example, the audiologist selects the audiometry test 202 (e.g., remotely from the clinician's computing device 180). The target word 204 can be selected from a plurality of target words for the audiometry test 202. The audiometry test 202 can be stored in various locations, such as locally on the computing device 110 or on a server 170. In the example, the audiometry test 202 is downloaded from server 170 to the computing device 110.
[0038] In the example, the audiometry test 202 can specify additional data. For instance, the audiometry test 202 can specify one or more complex factors to be added when providing the target word 204. For example, the audiometry test can add specific complex factors to the test audio in relation to specific conditions (e.g., white noise, Brownian noise, the sound of one or more conversations occurring, echo, reverberation, distortion, or other effects). The audiometry test 202 can further specify how the target word 204 is provided, such as using a specific volume. Furthermore, multiple different audio files containing the target word 204 (e.g., with different accents, pronunciations, or voice characteristics) can exist, and the audiometry test 202 can identify the specific audio file to be used for the target word 204. In another example, the target word 204 can be provided by a text-to-speech system, and the audiometry test 202 can specify which text-to-speech system to use or parameters to be used when generating speech (e.g., voice characteristics).
[0039] In some examples, method 200 can begin with operation 206.
[0040] Operation 206 includes selecting a target word 204 from the listening test 202. Selecting the target word 204 may include selecting either the first target word 204 of the listening test 202 or the next target word 204 of the listening test 202. This selection may be in a defined order (e.g., sequentially in a list) or may be arbitrary (e.g., randomly or pseudo-randomly selecting words until the total number of words has been used as part of the listening test 202). This selection may include selecting an identifier associated with the target word 204, an audio file associated with the target word 204, or the text of the target word 204. After operation 206, method 200 may move to operation 210.
[0041] Operation 210 includes providing test audio 212. Test audio 212 may include the target word 204 being pronounced. Providing test audio 212 may include causing sound to be generated based on the test audio. Test audio 212 may be provided using sound output component 114. For example, providing test audio 212 may include providing test audio 212 using a speaker, headset, or hearing device 120 to audibly generate test audio 212. Operation 210 may include obtaining test audio 212, for example, using an identifier of target word 204. For example, the identifier of target word 204 may be used to access a file storing test audio 212 or to stream test audio from server 170 (e.g., from audiometry data 172 stored thereon). In some examples, providing test audio 212 includes generating test audio 212. For example, computing device 110 may include text-to-speech functionality (e.g., using the AVSpeechUtterance category of iOS or the textToSpeech API of Android) or a remote computing device that provides text-to-speech functionality. The listening application 116 can provide the text of the target word 204 to the text-to-speech system to induce the production of audio output containing the target word 204. In other examples, someone can provide test audio 212, for example, by prompting with the target word 204. In some examples, the test audio 212 can be processed (e.g., to add noise or distortion). After operation 210, the flow of method 200 can move to operation 220.
[0042] Operation 220 includes obtaining a user response 222 from the user. The user response 222 can be obtained through a user interface. The user response 222 may include text data 224 or audio data 226. The user response 222 may be a user's response to the test audio 212. In some examples, before, during, or after providing the test audio 212, the listening application 116 may prompt the user to provide a user response 222 describing what the user perceives or understands from the test audio 212. Obtaining the user response 222 may include providing a user interface through which the user response 222 can be received. For example, the user interface may include a text input component 112, such as... Figure 1 As described above. Operation 220 may include making text input component 112 available for receiving text data 224 of user response 222. Operation 220 may include making audio input component (e.g., microphone) available for receiving audio data 226 of user response 222. In some examples, obtaining user response 222 may include checking user response 222. This check may include, for example, determining whether user response 222 includes valid words. If user response 222 includes invalid words (e.g., words not included in the user's language dictionary), computing device 110 may notify the user that the word is invalid (e.g., indicating that the word may be misspelled). In other examples, invalid words may be allowed as input. After operation 220, the flow of method 200 may move to operation 230.
[0043] Operation 230 includes converting the user response 222 into a user response speech representation 232. For example, if the user response 222 includes text data 224, this operation 230 may include performing a character-to-phoneme conversion on the text data 224 to generate the user response speech representation 232. The user response speech representation 232 is a representation of the user response 222 in speech form. For example, the user response speech representation 232 may be represented using the ARPABET or IPA speech system. ARPABET includes phonetic codes representing phonemes and allophones of General American English with different sequences of ASCII characters. IPA is a phonetic alphabet system primarily based on the Latin alphabet. Other representations or combinations of representations may be used.
[0044] In the example, one or more dictionaries can be used to generate the conversion. For example, one or more dictionaries can be searched based on text data 224 to obtain one or more user response speech representations 232. The one or more dictionaries may include, for example, a standard language-specific corpus dictionary (e.g., the CMU Speech Dictionary, which is the American English ARPABET dictionary) or an extended dictionary that can be (e.g., by an audiologist) adjusted to prolong or negate the dialogue (e.g., to correct dialect-specific pronunciation). In the example, a data structure (e.g., a hash table, database, or dictionary data structure) may allow lookups to be performed on the text input to produce one or more speech representations as output. As a supplement to or alternative to using one or more dictionaries, one or more rules can be used to perform the conversion. The rules may be symbol substitution rules that convert text data 224 into phonetic symbols. For example, the rules may be standard rules reflecting how written form is converted into spoken form in a specific language. In some examples, the conversion may be attempted first using the one or more dictionaries, and then, if one or more matches are not found in the one or more dictionaries, the conversion may be attempted using the rules.
[0045] When user response 222 includes audio data 226, converting user response 222 into a user response speech representation 232 may include transcribing the audio data 226 into one or more phonemes to generate the user response speech representation 232. For example, a speech-to-phoneme algorithm may be used. In other examples, the audio data 226 may be transcribed into text data (e.g., using a speech-to-text process), and then the text data may be converted into phonemes using the techniques described above.
[0046] After operation 230, the process of method 200 can be moved to operation 240.
[0047] Operation 240 includes obtaining a target word speech representation 242 of the target word 204. The target word speech representation 242 can be obtained in any of a variety of ways. For example, the target word speech representation 242 can be pre-generated and stored in association with the target word 204 as part of the listening test 202. In such an example, the target word speech representation 242 can be obtained by looking up the stored speech representation of the target word 204. In other examples, the target word 204 is generated via a similar process to obtaining the user response speech representation 232. For example, the textual form of the target word 204 can be used to obtain the target word speech representation 242 using a dictionary or rule-based method. After operation 240, the flow of method 200 can proceed to operation 250.
[0048] Operation 250 includes comparing the user response speech representation 232 and the target word speech representation 242 to obtain comparison data 252. In some examples, the comparison data 252 includes a score, and operation 250 may include determining the score based on the difference between the user response speech representation 232 and the target word speech representation 242. In some examples, an initial score 254 may be determined, and then the initial score may be modified to achieve a final score 256. In the presence of multiple user response speech representations 232 and / or target word speech representations 242, multiple different scores may be calculated between different pairs of these representations, and the highest score may be selected as the final score.
[0049] In the example, the initial score 254 can be based on whether specific parts of the speech representations 232 and 242 match. For example, in the case that the target word is a consonant-vowel-consonant (CVC) word, the initial score can be divided into three components of the form [c1, v, c2], where c1, v, c2 ∈ {0, 1} and a value of 1 indicates a match and a value of 0 indicates no match. Other kinds of configurations can be used and are not limited to the three-part form or the CVC form.
[0050] In the target word phonetic representation 242 is (“dog”) and the user's voice response indicates that... In the case of ("dig"), the initial score 254 can be [1,0,1], which reflects that the phonetic parts representing the consonants of the target word 204 match and the phonetic parts corresponding to the vowels do not match. In the example, inserting and replacing other phonemes will invalidate the correct phonemes. As a specific example, the phonetic representation of the target word 242 is (“sap”) and the user responded with a voice message indicating that 232 is (“snap”). Since the consonants and vowels of the target word speech representation 242 are in the user response speech representation 232, the initial score 254 can be [1,1,1]. However, since the additional phoneme “n” is added to the first consonant part, the initial score 254 can be modified to set the first consonant part incorrectly. This change can make the final score 256 [0,1,1].
[0051] In some examples, the comparison data 252 includes scores based on [c1, v, c2], such as a score s calculated as follows: s = c1 + v + c2 or The total score of the audiometry test 202 can be the sum, average, or another calculation result of some or all individual scores of the comparative data 252 associated with each target word 204.
[0052] In some examples, a score or other comparison data 252 can be generated based on the distance between the user response speech representation 232 and the target word speech representation 242. This difference can be determined based on any of a variety of techniques, such as Levenshtein distance, Hamming distance, Damerau–Levenshtein distance, or another distance technique. For example, in the case of using Levenshtein distance, the score or other comparison data could be the minimum number of deletions, insertions, and substitutions required to match the two representations 232 and 242. In some examples, a minimum cost calculation is added, which utilizes knowledge that a phoneme is a vowel or consonant to determine the most likely editing operation. In the presence of multiple phonemes (e.g., multiple different possible speech representations of the target word 204 or user response 222), all phonemes can be compared and the best match (e.g., the highest score) can be used. In some instances, the score or comparison data 252 can track or be based on changes in the use of consonants and vowels, for example, by using the following editing indicators:
[0053] Editing type abbreviation Correct consonants Cc Correct vowels Cv delete D Insert consonant Ic Insert vowel Iv Replace with the same phoneme type S Replace vowels with consonants Svc Replace consonants with vowels Scv
[0054] A correct consonant edit type can instruct the user to respond to a phonetic match between the consonant parts (e.g., "s" and "s") of the target word phonetic representation 242. A correct vowel edit type can instruct the user to respond to a phonetic match between the vowel parts (e.g., "ɑ" and "ɑ") of the target word phonetic representation 242. A deletion edit type can instruct the user to respond to a phonetic representation 232 that completely lacks components (e.g., consonants or vowels) of the target word phonetic representation. For example, in the case where the target word 204 is "dog", The user response voice representation of ("do") can be represented by 232. The second consonant part ("ɡ") of the target word phonetic representation 242 ("dog") is deleted. The Insert Consonant Edit type can instruct the user to add a consonant part that is not present in the target word 204 in response to phonetic representation 232. For example, in the case where the target word 204 is "do", The user-response speech representation 232 (“dog”) indicates the addition of a second consonant (“ɡ”) compared to the target word speech representation 242. The Insert Vowel Edit type can instruct the user-response speech representation 232 to add vowel components that are not present in the target word 204. For example, if the target word speech representation 242 is [d,ɑ,t] (“dot”) and the user-response speech representation 232 is... In the case of ("data"), The addition of a phoneme can be viewed as inserting a vowel edit. Replacing the edit type with the same phoneme type can instruct the user-response speech representation 232 to replace one phoneme with another phoneme of the same type. For example, in the point data example above, the user-response speech representation 232 contains... Instead of "ɑ", it can represent a substitution of the same phoneme type. The edit type from vowel to consonant can indicate where the target word's phonetic representation 242 includes vowels, and the user response phonetic representation 232 includes consonants. For example, in the target word's phonetic representation 242 is... ("ate") and the user's response voice indicates that In the case of (“kit”), replacing “e” with “k” indicates a change from a vowel to a consonant. The consonant-to-vowel edit type can instruct the user to respond that speech representation 232 includes a vowel where the target word speech representation 242 includes a consonant. For example, in the target word speech representation 242 is... (“kit”) and the user's response voice indicates that In the case of “ate”, replacing “k” with “e” indicates replacing a consonant with a vowel.
[0055] Figure 3 and Figure 4 An example of comparison data 252 obtained from operation 250 is shown. Figure 3 An example of the generation of a final score 256 based on a user response 222 to a target word 204 and other comparison data 252 is shown. In this example, the target word 204 is “tam”, which has a target word phonetic representation 242 of [t,ɑ,m]. The target word 204 is provided to the user, and a user response 222 of “term” is received, which has a user response phonetic representation 232 of [t,ε,r,m]. The comparison reveals the correct consonant (“t”), the same phoneme type substitution (replacing “ɑ” with the vowel “ε”), the inserted consonant (“r”), and the correct consonant (“m”). The comparison yields an initial score 254 of [1,0,1] because the first and last consonants match and the vowel does not match. Further processing yields a final score of [1,0,0] because the last consonant is incorrect because “r” was inserted in addition to the correct “m”. Figure 4Another example of generation is shown based on the final score 256 of the user response 222 to the target word 204 and other comparison data 252. In this example, the target word 204 is “hen”, which has a target word phonetic representation 242 of [h, ε, n]. The target word 204 is provided to the user, and a user response 222 of “end” is received, which has a user response phonetic representation 232 of [ε, n, d]. The comparison reveals the deletion of the consonant (“h”), the correct vowel (“ε”), the correct consonant (“n”), and the insertion of the consonant (“d”). The comparison yields an initial score 254 of [0, 1, 1] because the vowel and the last consonant match. Further processing yields a final score 256 of [0, 1, 0] because the last consonant is incorrect, the reason being that an incorrect consonant (“d”) was inserted as part of the final consonant in addition to the correct consonant.
[0056] return Figure 2 As can be seen above, any of the various techniques can be used to compare the user response speech representation 232 and the target word speech representation 242 to generate comparison data 252, and the comparison data 252 can include any of the various types of data. After operation 250, the flow of method 200 can move to operation 260. In some examples, if an additional target word 204 is present in the audiometry test 202, the flow of method 200 can return to operation 206 to select the next target word 204. If no additional target word 204 is to be provided, the flow of the method can move to operation 260.
[0057] Operation 260 includes determining a user's hearing ability 262 based on comparison data 252. The user's hearing ability 262 can take any of a variety of forms. In some instances, the user's hearing ability 262 can be a qualitative score, such as the sum or average of one or more scores determined in the comparison data 252. Alternatively or additionally, the determined hearing ability 262 can include the recognition of specific sounds, frequencies, phonemes, syllables, speech segments, or other aspects that the user has difficulty hearing or understanding. Alternatively or additionally, hearing ability 262 can involve the ability to listen using a specific ear or under specific conditions. For example, audiometry tests can add specific complicating factors to test audio in relation to specific conditions (e.g., white noise, Brownian noise, the sound of one or more conversations occurring, echoes, reverberation, distortion, or other effects), and hearing ability 262 can also include data related to conditions that the user hears better or worse.
[0058] In some examples, the determination of the user's hearing ability 262 is further based on the amount of time it takes for the user to respond to the test audio 212. For example, method 200 may include determining the amount of time between the end of the test audio 212 and the start of receiving the user's response 222. The amount of time may be an indication of how easy or difficult the user perceives the test audio 212 to be.
[0059] After operation 260, the process of method 200 can be moved to operation 270.
[0060] Operation 270 includes performing treatment measures 272 based on hearing ability 262. For example, treatment measures may include providing the user with hearing aids or modifying the user's existing hearing aids (e.g., adjusting the low-frequency gain of the device in response to vowel errors or adjusting the high-frequency gain in response to fricative errors). In some examples, treatment measures 272 include diagnosing the user with a specific hearing condition.
[0061] In the example, treatment measure 272 is a measure involving the treatment of a medical condition associated with the recipient's auditory system. Various treatment measures 272 may be identified or recommended. In the example, treatment measure 272 includes, for example, reporting the performance quality of the user's hearing to a clinician (e.g., to help guide treatment) or a caregiver (e.g., to help assure the caregiver that the user's hearing device is working as intended). In the example, treatment measure 272 includes providing a measure of the recipient's ability to perceive a particular sound. In the example, treatment measure 272 includes recommending corrective measures.
[0062] In examples, treatment measure 272 includes recommended corrective measures (e.g., reconfiguration, reprogramming, or modification of treatment, such as by progressing from a unilateral prosthesis to a bilateral prosthesis). In some examples, treatment measure 272 may include a recommendation or modification of one or more auditory device settings 122 of the user's auditory device 120. Modifying auditory device settings 122 can cause a modification to the ongoing treatment provided by auditory device 120. Based on audiometry test results 500, it can be determined that auditory device settings 122 suboptimally cause auditory perception in the recipient relative to a particular auditory input (e.g., a particular phoneme or sound), and that one or more changes to auditory device settings 122 may improve the performance of auditory device 120. Based on this determination, information relating to said one or more changes can be provided to the recipient, caregiver, or clinician (e.g., by reporting device 180 to a clinician). In some examples, said one or more changes are automatically adopted by auditory device 120 itself. The auditory device settings 122 are then changed, thereby modifying the ongoing operation of auditory device 120. In some examples, scene-specific auditory device settings 122 are changed. For example, auditory device settings 122 associated with a voice mode (e.g., as determined by a scene classifier of auditory device 120) are changed, but auditory device settings are not changed in other modes (e.g., music or wind modes).
[0063] Treatment measures 242 may include using hearing ability 262 or audiometry test results 500 as input to a hearing rehabilitation training plan. The hearing rehabilitation plan can take any of a variety of forms. For example, a hearing rehabilitation plan may include the use of rehabilitation exercises, such as a playlist of words containing sounds that the user cannot hear clearly. As a specific example, audiometry application 116 may (automatically or based on clinician feedback) provide the user with specific exercises to perform, such as listening to a list of words beginning with a specific phoneme (e.g., a specific phoneme indicated by the hearing test that the recipient has difficulty understanding).
[0064] Example audiometry test results
[0065] Figure 5 is from Figure 5A , Figure 5B and Figure 5C The composition illustrates an example audiometry test result 500 of audiometry test 202 (e.g., performed using method 200). Figure 5A As shown, the audiometry test results 500 may include a results table 501, phoneme analysis 590, and one or more recommended treatment measures 272. The audiometry test results 500 may be generated, for example, as part of operations 250 and / or 260 as described above.
[0066] Figure 5BAn example implementation of the results table 501 is shown. Results table 501 can be a data table generated after providing audiometry test 202. Results table 501 can be provided to clinicians or analytical procedures to determine a recipient's hearing ability. Results table 501 includes ten entries 502, each corresponding to a different word response period. Results table 501 also includes several fields, including a file field 510, a target word field 520, a user response field 530, a target word phoneme field 540, a user response phoneme field 550, a phoneme score field 560, a word score field 570, and a response time field 580. File field 510 is a field describing which file (e.g., an audio file) was used to generate the audio containing target word 204. Target word field 520 describes the text of target word 204. User response field 530 describes user response 222 in text form. Target word phoneme field 540 is a field describing the phonetic representation 242 of the target word. User response phoneme field 550 is a field describing the phonetic representation 532 of the user response. As shown in the figure, more than one distinct phoneme representation can be stored in association with phoneme fields 540 and 550. Phoneme score field 560 is a field describing phoneme scores. In the example shown, the phoneme score is in the three-part CVC form described above: [c1,v,c2], where c1,v,c2∈{0,1}. Other configurations are also possible. Word score field 570 describes word scores. In the example shown, the word score is the count of matches in phoneme score field 560. Response time field 580 is a field describing the amount of time between providing the target word 204 and the user beginning to provide a user response. In the example shown, response time field 580 describes the data in milliseconds.
[0067] Figure 5C An example implementation of phoneme analysis data 590 is shown. In the example shown, phoneme analysis data 590 includes phoneme result data 592 and score data 594.
[0068] Phoneme result data 592 may include phoneme-specific data relating to the hearing test 202. Phoneme result data 592 may be expressed in any of a variety of ways. In the example shown, phoneme result data 592 is implemented as a table showing the degree to which the user's response was correct or incorrect relative to a particular phoneme for each phoneme presented during the test. Such data can be used, for example, to help identify specific sounds that the user has difficulty hearing. Some phonemes may tend to be produced at higher or lower frequencies than others. For example, the phoneme "IY" in the word "bean" is at a relatively higher frequency than the phoneme "OW" in the word "hope". A high incidence of incorrect answers related to the phoneme "IY" may indicate that the user has difficulty hearing high-frequency sounds. Conversely, a high incidence of incorrect answers related to phonemes associated with low-frequency sounds may indicate that the user has difficulty hearing low-frequency sounds.
[0069] Score data 594 may include data related to the user's score in the listening test 202. For example, as shown, score data 594 may include a total phoneme score in the form of the total number of correct answers for CVC words in the form of [c1,v,c2]. Alternatively, score data 594 may include a total word score (e.g., the total number of correct components) or a percentage of correct answers. Alternatively, other data may be used.
[0070] Hearing devices
[0071] The audiometry test described herein can be used with any user of a variety of types, including users who are recipients of one or more hearing devices 120. Hearing devices 120 may include devices related to the recipient's hearing. Hearing devices 120 may take various forms, including cochlear implants, electroacoustic devices, percutaneous bone conduction devices, passive percutaneous bone conduction devices, active percutaneous bone conduction devices, middle ear devices, fully implantable hearing devices, mostly implantable hearing devices, auditory brainstem implants, hearing aids, dental-anchored hearing devices, personal sound amplification products, other auditory prostheses, and combinations of the above devices (e.g., binaural systems, which include a prosthesis for the recipient's first ear and the same or different types of prostheses for the second ear).
[0072] The audiometry test described herein can be used to determine the operational status of a user’s hearing device 120 and whether one or more changes to the hearing device settings 122 may be beneficial.
[0073] Example implementations of hearing devices in Figure 6 (showing a cochlear implant) and Figure 7 (A more detailed description is provided in the section showing the percutaneous bone conduction device.)
[0074] Hearing devices—cochlear implant system
[0075] Figure 6 An example cochlear implant system 610 is shown that can benefit from the use of the technology disclosed herein. The cochlear implant system 610 includes an implantable component 644, which typically has an internal receiver / transceiver unit 632, a stimulator unit 620, and elongated leads 618. The internal receiver / transceiver unit 632 allows the cochlear implant system 610 to receive signals from and / or transmit signals to an external device 650. The external device 650 may be a head-worn button sound processor that includes a receiver / transceiver coil 630 and a sound processing component. Alternatively, the external device 650 may simply be a transmitter / transceiver coil communicating with a behind-the-ear device that includes the sound processing component and a microphone.
[0076] The implantable component 644 includes an internal coil 636 and a magnet (not shown) preferably fixed relative to the internal coil 636. The magnet may be embedded together with the internal coil 636 in a flexible silicone or other biocompatible sealant. The transmitted signal typically corresponds to an external sound 613. The internal receiver / transceiver unit 632 and the stimulator unit 620 are hermetically sealed within a biocompatible housing, and they are sometimes collectively referred to as the stimulator / receiver unit. The included magnet (not shown) facilitates operational alignment of the external coil 630 and the internal coil 636, enabling the internal coil 636 to receive electrical and stimulation data from the external coil 630. The external coil 630 is contained within an external portion. An elongated lead 618 has a proximal end connected to the stimulator unit 620 and a distal end 646 implanted in the cochlea 640 of the recipient. The elongated lead 618 extends from the stimulator unit 620 through the mastoid bone 619 of the recipient to the cochlea 640. The thin lead 618 is used to deliver electrical stimulation to the cochlea 640 based on stimulation data. Stimulation data can be created based on external sound 613 using a sound processing component and based on auditory device settings 122.
[0077] In some examples, the outer coil 630 transmits electrical signals (i.e., power and stimulation data) to the inner coil 636 via a radio frequency (RF) link. The inner coil 636 is typically a wire antenna coil with multiple turns of electrically insulated single or multiple strands of platinum or gold wire. The electrical insulation of the inner coil 636 can be provided by a flexible silicone molding. Various types of energy transfer, such as infrared (IR), electromagnetic, capacitive, and inductive transfer, can be used to transfer power and / or data from an external device to the cochlear implant. While the above description has described the inner and outer coils as being formed of insulated wire, in many cases, the inner and / or outer coils can be implemented via conductive traces.
[0078] Hearing devices—transcutaneous bone conduction devices
[0079] Figure 7 This is a view of an example percutaneous bone conduction device 700 from which the use of the techniques disclosed herein may benefit. For example, one or more aspects of the disclosed techniques can be used to customize the sensory prosthetic setting 146 of the device 700. The bone conduction device 700 is positioned behind the outer ear 701 of the recipient of the device. The bone conduction device 700 includes a sound input element 726 for receiving sound signals 707. The sound input element 726 may be a microphone, a pickup coil, or a similar element. In this example, the sound input element 726 may be located, for example, on or within the bone conduction device 700, or on a cable extending from the bone conduction device 700. Additionally, the bone conduction device 700 includes a sound processor (not shown), a vibrating electromagnetic actuator, and / or various other operating components.
[0080] More specifically, the sound input element 726 converts the received sound signal into an electrical signal. These electrical signals are processed by a sound processor. The sound processor generates a control signal that causes the actuator to vibrate. In other words, the actuator converts the electrical signal into a mechanical force to apply vibration to the recipient's skull 736. The conversion of the electrical signal into a mechanical force can be based on the sensory prosthesis setting 146, such that different sensory prosthesis settings 146 can cause different mechanical forces to be generated from the same sound signal 707.
[0081] The bone conduction device 700 also includes a coupling device 740 for attaching the bone conduction device 700 to a recipient. In the example shown, the coupling device 740 is attached to an anchoring system (not shown) implanted in the recipient's body. An exemplary anchoring system (also referred to as a fixation system) may include a percutaneous abutment fixed to a skull 736. The abutment extends from the skull 736 through muscle 734, fat 728, and skin 732, allowing the coupling device 740 to be attached thereto. This percutaneous abutment provides the coupling device 740 with an attachment site that facilitates efficient transmission of mechanical forces.
[0082] Example computing system
[0083] Figure 8 Examples of suitable computing systems 800 that can be used to implement one or more of the disclosed examples are shown. Computing systems, environments, or configurations suitable for use with the examples described herein include, but are not limited to, personal computers, server computers, handheld devices, laptop devices, multiprocessor systems, microprocessor-based systems, programmable consumer electronics (e.g., smartphones), network PCs, minicomputers, mainframe computers, tablet computers, distributed computing environments including any of the aforementioned systems or devices, and so on. Computing system 800 may be a single virtual or physical device operating in a networked environment via a communication link to one or more remote devices. Remote devices may be auditory devices (e.g., auditory device 120), computing devices 110, personal computers, servers, routers, networked personal computers, peer-to-peer devices, or other public network nodes. In the examples, computing devices 110 and server 170 include one or more components of computing system 800 or variations thereof. Furthermore, in some examples, auditory device 120 includes one or more components of computing system 800.
[0084] In the basic configuration, the computing system 800 includes one or more processors 802 and memory 804.
[0085] The one or more processors 802 may include one or more hardware or software processors (e.g., central processing units or microprocessors) capable of receiving and executing instructions. The one or more processors 802 may communicate with and control the performance of other components of the computing system 800.
[0086] Memory 804 may include one or more computer-readable storage media based on software or hardware, operable to store information accessible by the one or more processors 802. Among other things, memory 804 may store instructions and other data that can be executed by the one or more processors 802 to implement an application or cause the operations described herein to be performed. Memory 804 may be volatile memory (e.g., RAM), non-volatile memory (e.g., ROM), or a combination thereof. Memory 804 may include temporary or non-temporary memory. Memory 804 may also include one or more removable or non-removable storage devices. In examples, memory 804 may include RAM, ROM, EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory, optical disc storage, magnetic storage, solid-state storage, or any other memory medium that can be used to store information for later access. In examples, memory 804 encompasses modulated data signals (e.g., signals whose one or more characteristics are set or changed in a manner such as encoding information in the signal), such as carrier waves or other transmission mechanisms, and includes any information delivery medium. By way of example and not limitation, memory 804 may include wired media, such as wired networks or direct wiring connections, and wireless media, such as acoustic, RF, infrared and other wireless media or combinations thereof.
[0087] In the example shown, system 800 also includes a network adapter 806, one or more input devices 808, and one or more output devices 810. System 800 may include other components such as a system bus, component interfaces, a graphics system, a power supply (e.g., a battery), and other components.
[0088] Network adapter 806 is a component of computing system 800 that provides network access. Network adapter 806 can provide wired or wireless network access and can support one or more of various communication technologies and protocols, such as Ethernet, cellular, Bluetooth, near field communication, and RF (radio frequency), etc. Network adapter 806 may include one or more antennas and associated components configured for wireless communication according to one or more wireless communication technologies and protocols.
[0089] The one or more input devices 808 are devices through which the computing system 800 receives input from a user. The one or more input devices 808 may include physically actuated user interface elements (e.g., buttons, switches, or dial pads), touchscreens, keyboards, mice, pens, and voice input devices, as well as other input devices.
[0090] The one or more output devices 810 are devices through which the computing system 800 can provide output to the user. The output devices 810 may include displays, speakers, printers, and other output devices.
[0091] ***
[0092] It should be understood that while specific uses of this technology have been described and discussed above, the disclosed technology can be used with a variety of devices based on many examples of this technology. The foregoing discussion is not intended to suggest that the disclosed technology is only suitable for implementation in systems similar to those shown in the accompanying drawings. In general, additional configurations can be used to practice the processes and systems described herein, and / or some aspects can be excluded without departing from the processes and systems disclosed herein.
[0093] This disclosure describes some aspects of the invention with reference to the accompanying drawings, which illustrate only some possible aspects. However, other aspects may be embodied in many different forms and should not be construed as limited to those set forth herein. Rather, these aspects are provided to make this disclosure exhaustive and complete and to fully convey the scope of possible aspects to those skilled in the art.
[0094] It should be understood that this document is not intended to limit the system and process to the specific aspects described with respect to the accompanying drawings (e.g., parts, components, etc.). Therefore, additional configurations can be used to practice the methods and systems described herein, and / or some aspects described may be excluded without departing from the methods and systems disclosed herein.
[0095] Similarly, where the steps of a method are disclosed, these steps are described for the purpose of illustrating the methods and systems of the present invention and are not intended to limit this disclosure to a specific sequence of steps. For example, these steps may be performed in a different order, two or more steps may be performed simultaneously, additional steps may be performed, and the disclosed steps may be excluded without departing from this disclosure. Furthermore, the disclosed process may be repeated.
[0096] Although specific aspects have been described herein, the scope of this technology is not limited to those specific aspects. Those skilled in the art will recognize other aspects or modifications within the scope of this invention. Therefore, specific structures, operations, or media are disclosed only as illustrative aspects. The scope of this technology is defined by the following claims and any of their equivalents.
Claims
1. A computer-readable medium storing instructions, which, when executed by one or more processors, cause the one or more processors to: Select (206) target words (204) from multiple words in the audiometry test (202); Provide test audio (212) of the target word (204) pronounced by (210); User responses (222) are obtained (220) as text data (224) through the user interface; The text data (224) is converted (230) into a user-response speech representation (232), wherein the conversion of the text data (224) into the user-response speech representation (232) includes: Perform a character-to-phoneme conversion on the text data (224) to generate the user response speech representation (232); Obtain the target word speech representation (242) of the target word (204) mentioned in (240); Compare (250) the user response speech representation (232) and the target word speech representation (242) to obtain comparison data (252); and The user's hearing ability is determined (260) based on the comparison data.
2. The computer-readable medium of claim 1, wherein comparing (250) the user response speech representation (232) and the target word speech representation (242) to obtain the comparison data (252) comprises: The score (254, 256) is determined based on the difference between the user response speech representation (232) and the target word speech representation (242).
3. The computer-readable medium of claim 2, wherein comparing the user response speech representation (232) and the target word speech representation (242) to obtain the comparison data (252) comprises: The difference between the user response speech representation (232) and the target word speech representation (242) is determined based on the Levenshtein distance between the user response speech representation (232) and the target word speech representation (242).
4. The computer-readable medium as described in any one of claims 1 to 3, The comparison (250) between the user response speech representation (232) and the target word speech representation (242) includes: Determine the initial score (254) and modify the initial score (254) to obtain the final score (256); The target words (204) mentioned therein are consonant-vowel-consonant words; The comparison (250) between the user response speech representation (232) and the target word speech representation (242) includes determining the number and type of edits required to convert the user response speech representation (232) into the target word speech representation (242) or vice versa; The comparison (250) between the user response speech representation (232) and the target word speech representation (242) includes determining phoneme scores; The computer-readable medium mentioned therein is a non-transitory computer-readable medium; The computer-readable medium described therein is a component of a telephone or tablet computer; or The instructions are stored as a downloadable package.
5. A method comprising: Provide the user (210) with test audio (212) including the target word (204) being pronounced; Receive (220) a user response (222) including text data (224) from the user; The user response (222) is converted (230) into a user response speech representation (232), wherein the conversion of the user response (222) into a user response speech representation (232) includes: performing a character-phoneme conversion on the text data (224) to generate the user response speech representation (232); The target word (204) is converted (240) into a speech representation of the target word (242); Compare (250) the user response speech representation (232) and the target word speech representation (242) to obtain comparison data (252); and The user's hearing ability is determined (260) based on the comparison data (252).
6. The method of claim 5, wherein determining the user's hearing ability (262) comprises: The specific phonemes that the user has difficulty hearing are determined based on the comparison data (252).
7. The method of claim 5 or 6, wherein obtaining the target word speech representation (242) of the target word (204) comprises: Obtain multiple speech representations of the target word (204); and Select the speech representation that best matches the user response speech representation from the plurality of speech representations as the target word speech representation (242).
8. The method of claim 5 or 6, wherein comparing (250) the user response speech representation (232) and the target word speech representation (242) to obtain the comparison data (252) comprises: The score is determined based on the difference between the user response speech representation (232) and the target word speech representation (242).
9. The method of claim 8, wherein the difference is the Levenshtein distance between the user response speech representation (232) and the target word speech representation (242).
10. The method of claim 5 or 6, further comprising: Determine the amount of time between the end of the test audio (212) and the start of receiving the user response (222). The determination (260) of the user's hearing ability is further based on the amount of time.
11. The method as described in claim 5 or 6, The method further includes determining the total number of incorrect phonemes; The comparison between the user response speech representation (232) and the target word speech representation (242) includes: Determine the initial score (254) and modify the initial score (254) to obtain the final score (256); The target words (204) mentioned therein are consonant-vowel-consonant words; The comparison of the user response speech representation (232) and the target word speech representation (242) includes determining the number and type of edits required to convert the user response speech representation (232) into the target word speech representation (242) or vice versa; The comparison of the user response speech representation (232) and the target word speech representation (242) includes determining phoneme scores; Obtaining the user response voice representation (232) includes: obtaining multiple voice representations of the user response (222); And select the speech representation that best matches the speech representation of the target word (242) from the plurality of speech representations as the user response speech representation (232); The target word (204) is the first word among multiple test words; The determination of the user's hearing ability based on the comparison data (252) is performed without the need for a clinician; or The method described herein is performed by a telephone or tablet computer.
12. A system (100) comprising: Text input component (112); Sound output component (114); One or more processors (802); as well as A memory (804) stores instructions that, when executed by the one or more processors, cause the one or more processors to: Test audio (212) including the pronounced target word (204) is provided (210) via the sound output component (114); Text data (224) is obtained from the text input component (220); The text data (224) is converted into a user-response speech representation (232), wherein converting the text data (224) into the user-response speech representation (232) includes: performing a character-phoneme conversion on the text data (224) to generate the user-response speech representation (232); Obtain the target word speech representation (242) of the target word (204); Compare the user response speech representation (232) and the target word speech representation (242) to obtain comparison data (252); and The user's hearing ability is determined based on the comparison data (252) (262).
13. The system of claim 12, further comprising: A first computing device, the first computing device comprising: The sound output component (114); The text input component (112); The one or more processors (802); and The memory (804), The memory (804) further stores instructions that, when executed, cause the one or more processors (802) to: The test audio is obtained from a second computing device located away from the first computing device.
14. The system as described in claim 13, The first computing device is a consumer computing device, and the second computing device is a server.
15. The system as described in claim 13 or 14, The sound output component (114) mentioned above is a loudspeaker; The audio output component (114) is wirelessly connected to the one or more processors; The text input component (112) mentioned above is a physical keyboard; The text input component (112) mentioned above is a virtual keyboard; or The system further includes an auditory device (120) having the sound output component (114).
Citation Information
Patent Citations
Auditory signal processing
US9338567B2
Pre-processing of a channelized music signal
US9473852B2
Mobile phone capable of performing hearing test and hearing test method through the same
KR1020060097404A
Individualized rehabilitation training of a hearing prosthesis recipient
US20180012511A1
Computer implemented method and apparatus for recognition of speech patterns and feedback
WO2019215459A1