Methods and apparatus for hearing training

The method enhances hearing training by using binaural and spatialized audio to mimic real-life scenarios, offering personalized and effective training for users with varying hearing loss levels.

JP7853307B2Active Publication Date: 2026-04-28EARGYM LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
EARGYM LTD
Filing Date
2022-04-27
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

Existing hearing training methods fail to accurately reflect real-life situations, are inaccessible to many users, and lack flexibility, often requiring the same regime for users with varying hearing loss levels.

Method used

A computer-implemented hearing training method using a user device with binaural audio and spatialized audio, where the user distinguishes overlapping audio signals and receives feedback, mimicking real-life scenarios, and adjusting difficulty based on user performance.

Benefits of technology

Improves listening skills, including sound detection, location, and clarity in noisy environments, by providing realistic training scenarios and adapting to individual user needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007853307000001
    Figure 0007853307000001
  • Figure 0007853307000002
    Figure 0007853307000002
  • Figure 0007853307000003
    Figure 0007853307000003
Patent Text Reader

Abstract

A computer-implemented method, user device, and non-transitory computer-readable medium having instructions stored thereon for performing hearing training using a user device having a user interface and an audio output, the training comprising: providing, using the audio output, a background audio signal and a target audio signal, the target audio signal at least partially overlapping with the background audio signal, the target audio signal defining information to be determined by a user, one or both of the background audio signal and the target audio signal including binaural audio; receiving, at the user interface, user input corresponding to a user evaluation of the information defined by the target audio signal; and providing feedback to the user based on the user evaluation indicated by the user input.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a computer-implemented method for performing hearing training, particularly a method using a smart user device. In particular, the present invention provides a particularly effective and convenient means for training a user's hearing.

Background Art

[0002] Humans hear and recognize sounds by detecting air vibrations through their ears. Hearing is an important way for humans to interact with the environment.

[0003] However, hearing loss is a common disease for many people. Age-related hearing loss occurs gradually over time and is particularly common in people over 60 years old. Noise-induced hearing loss can also occur when a person is exposed to loud noises such as machinery, explosions, gunshots, or loud music.

[0004] It will be understood that people suffering from hearing loss experience both a decrease in sensory response to sound and an increase in cognitive load because their brains struggle to adapt to changes in hearing and have to work harder to process and distinguish noise. Therefore, there is a great need for training methods that enable users to better recognize sounds and / or reduce the cognitive load associated with hearing. In fact, there is a particular shortage of tools that can be used to help people with age-related hearing loss or noise-related hearing loss.

[0005] Methods for measuring, monitoring, and training hearing have been proposed heretofore. However, these approaches typically have many common problems.

[0006] Most importantly, existing approaches do not accurately reflect real life. The tasks and techniques involved provide insufficient training in users' daily lives. Furthermore, testing and training are often conducted in laboratories and require specialized equipment. As a result, these approaches may be inaccessible to many users. Finally, although different users may have vastly different hearing loss ranges and levels, conventional methods often require the same regime for each user. This lack of flexibility reduces the effectiveness of these existing approaches. [Overview of the Initiative] [Problems that the invention aims to solve]

[0007] Therefore, there is a clear need for improved hearing training methods, systems, and equipment that overcome at least some of the problems identified above. [Means for solving the problem]

[0008] According to one aspect of the present invention, a computer implementation method for performing hearing training using a user device having a user interface and audio output, comprising: providing a background audio signal and a target audio signal using the audio output, wherein the target audio signal at least partially overlaps with the background audio signal, and the target audio signal defines information to be determined by the user; and where one or both of the background audio signal and the target audio signal include binaural audio, and the user interface receives user input corresponding to the user evaluation of the information defined by the target audio signal, and provides feedback to the user based on the user evaluation indicated by the user input.

[0009] This aspect of the present invention will be understood to provide a particularly practical and therefore effective method for training a user's hearing. The training involves mimicking sounds and situations that the user experiences in everyday life.

[0010] The user must distinguish the target audio signal from the background audio signal, determine or identify the information provided by the target audio, and provide user input related to understanding the target audio. Upon receiving feedback, the user can recognize whether their evaluation of the target audio signal was correct, thereby improving their listening skills. Listening skills that can be improved through this method include sound detection, location, distinction, clarity in quiet environments, and clarity in noisy environments.

[0011] This invention utilizes binaural audio, which means audio containing two distinct audio channels (i.e., right and left audio channels), each configured to independently deliver the same sound to each of the listener's ears, and the audio channels differ based on the assumed relative placement of the ears. For example, the binaural audio described above may be recorded using two microphones positioned on either side of a model head or a human head in a manner similar to human ears. This approach is often referred to as "binaural recording" or the binaural recording process. The difference in sound recorded by each microphone, i.e., the difference between the two audio channels of the binaural audio, is defined by the relative positions of the two microphones on a model or human head that ideally approximates the relative positions of the listener's ears. Alternatively, the binaural audio described above may be artificially synthesized or generated from conventional audio by forming two audio channels from conventional audio using a transfer function (i.e., HRTF (Head-Related Transfer Function)) that defines the assumed relationship between the user's ears. In particular, the difference between the right and left audio channels of binaural audio may be based on the Head-Related Transfer Function (HRTF), which characterizes how each ear receives sound from a specific point in space. In a preferred example, the binaural audio may be adapted or modified according to the relative position and orientation between the user's head and the apparent sound source in the binaural audio. Various software products are available that generate binaural audio from mono or stereo sound, such as the "AMBEO Orbit" plugin from Sennheiser Electronic GmbH & Co. and the "DearVR MICRO" plugin from Dear Reality GmbH, both of which are plugins for the audio production software "Pro Tools" manufactured by Avid Technology, Inc.

[0012] In summary, binaural audio is a concrete example of stereo audio (also referred to as "stereo audio" or "stereo"), where the difference between the right and left audio channels is based on a hypothetical relationship between the listener's ears. This hypothetical relationship corresponds to the Head-Related Transfer Function (HRTF) of each ear, which does not exist in conventional stereo audio.

[0013] As described above, binaural audio in the target audio signal and / or background audio signal may be binaurally recorded and / or generated or synthesized from their respective input audio signals (e.g., their respective mono or stereo input audio signals).

[0014] The inventors have recognized the particular benefits of using binaural audio in hearing training. In contrast to conventional monaural audio (also referred to as "monaural audio," "mono audio," or "mono") or stereo audio, binaural audio is particularly similar to the sounds a user hears in real life. This allows for training that particularly accurately mimics situations and tasks that users commonly experience. Furthermore, binaural audio allows for training users in spatial resolution of sound—that is, their ability and skill to identify the location and source of sounds based on the differences in how sounds are heard in each ear. This is a skill commonly used by healthy individuals but is particularly difficult for people with hearing loss, especially those with significant differences in hearing between their two ears. Training such users in spatial resolution of sound is particularly difficult using conventional monaural or stereo sound.

[0015] Furthermore, the use of binaural audio enables spatialization. Spatialization is the process of modifying an audio signal so that the listener can localize the audio signal, making the sound defined by the audio signal appear to originate from a specific location. Spatialized audio makes it possible to create particularly complex listening training situations. Therefore, all binaural audio described herein may be spatialized audio, and references to binaural audio may be replaced with spatialized audio as needed.

[0016] In a preferred example, the method includes tracking the position and orientation of the user's head relative to the apparent sound source in the binaural audio of the target audio signal and / or background audio signal. The method may then include adapting the binaural audio in the target audio signal and / or background audio signal based on the position and / or orientation of the user's head relative to the apparent sound source, so that the position of the apparent sound source appears to the user to be constant. Thus, the method may include so-called “head tracking.” Head tracking may be performed using a camera configured to detect the movement of the user’s head, an accelerometer or sensor mounted on the head of the user or user device, and / or any other suitable method. Alternatively, the position and orientation of the user’s head may be determined or assumed based on the position and / or orientation of the user device. The binaural audio produced through these steps is also called “reactive binaural audio” or “adaptive binaural audio with head tracking.” Such binaural audio may be generated from conventional monaural or audio signals using a Head-Related Transfer Function (HRTF) that depends on the listener's position and orientation. Suitable tools for achieving this include the plugins mentioned above.

[0017] The use of binaural audio, which depends on the user's position and orientation relative to the apparent sound source, is particularly practical. Therefore, hearing training is especially effective in improving the hearing skills that users use in everyday life.

[0018] Spatialized audio offers particular advantages when combined with VR (Virtual Reality) settings and training environments, as will be further explained below.

[0019] The target audio signal and the background audio signal overlap at least partially. This ensures that the signals are delivered to the user simultaneously or synchronously, so that at least some sounds within each signal are heard together by the user. Users with hearing impairments may have difficulty distinguishing such complex arrangements of overlapping audio signals. Such overlapping audio signals accurately reflect real life. Therefore, hearing training according to this aspect of the present invention is particularly effective. In particular, background audio signals can act as a distraction to the user, making it more difficult for the user to identify and interpret the target audio signal amid the background audio "noise."

[0020] Preferably, the background audio signal includes binaural audio, and the background audio signal defines two or more sounds that have different apparent sound sources to the user. The use of a background audio signal that includes multiple sounds positioned at different apparent locations is particularly practical. In addition, or instead, the background audio signal and the target audio signal each include binaural audio, and the background audio signal defines one or more sounds that have different apparent sound sources to the user, distinct from the sound defined by the target audio signal. By providing multiple sounds positioned at different apparent sound sources, it is possible to generate a particularly realistic "soundscape" (similar to a landscape) that corresponds to situations the user might encounter in everyday life. Therefore, hearing training is particularly effective in improving the user's hearing in normal situations. The binaural audio for each sound can be adjusted based on the user's position and orientation relative to the apparent sound source in the manner described above.

[0021] By providing feedback to the user based on the results of their user evaluation, the user can recognize whether or not they have correctly identified the information provided by the target audio signal. This encourages the user to improve their listening skills. For example, the method may include providing feedback to the user in the form of audio instructions, visual instructions, and / or tactile instructions (e.g., vibration of the user's device) to indicate whether or not they have correctly understood the information defined in the target audio signal. In further examples, the feedback may be based in part on the time it takes for the user to evaluate the target audio signal and provide user input.

[0022] As described above, the target audio signal defines information to be determined by the user. The target audio signal may be configured to transmit or provide information in any suitable manner. For example, the information defined by the target audio signal may include the target location in the training environment, the content of the target audio signal, preferably the linguistic content of the speech within the target audio signal, and / or the similarity and / or relationship with a second target audio signal.

[0023] Therefore, the target audio signal may provide information directly to the user through the content of the signal, for example, the information being defined by specific sounds and / or words within the target audio signal. Alternatively, the target audio signal may provide information indirectly to the user, in which case the user may be required to interpret the target audio signal in order to identify the information (e.g., a specific location, and / or similarity or relationship to other sounds).

[0024] The actions or tasks required of the user during hearing training may take various forms, depending on how the information is conveyed to the user by the target audio signal.

[0025] For example, when the information defined by the target audio signal includes the target position within the training environment, providing the target audio signal involves receiving, at the user interface, one or more preliminary user inputs respectively corresponding to intermediate positions within the training environment, and changing one or more characteristics of the target audio signal based on the relative position between the intermediate position and the target position within the training environment.

[0026] Depending on the distance and other relationships between the target position and the intermediate position, the sound heard by the user can change. Thus, the user may identify the target position based on the change in the target audio signal and / or the background audio signal when the input intermediate position changes. Thus, the method may include comparing the intermediate position input by the user with the target position and changing the audio signal provided to the user based on the result of this comparison. Preferably, this comparison is performed by the user device, but this is not essential and may be performed by a separate device or system. After identifying what the user believes to be the target position, the user may provide a user input corresponding to the evaluation of the target position. Thus, the user input received during the method may be an indication of the position that the user believes corresponds to the target position based on the change in the target audio signal. When the user input corresponds to the target position and / or is within a predetermined distance from the target position, the user may be understood to have accurately identified the position of the target position.

[0027] Thus, in this training method, the user searches (seek) and forages for the target position within the training environment that is visually hidden from the user based on the provided audio signal. Thereby, the user's skills in sound detection, identification, and position determination can be improved. The training environment may be the physical environment where the user is located, but more preferably, it is the training environment presented to the user (further described below).

[0028] These seeking or foraging training methods are designed to improve the user's hearing skills in sound detection. Cognitively, the user can improve their attention duration and spatial working memory. Thus, these types of methods are designed to help the user ensure safety in daily life and improve their ability to perform tasks such as spatial activities and sports.

[0029] Changing one or more properties of a target audio signal may include changing the volume of the target audio signal relative to the background audio signal, changing the content of the target audio signal, changing the pitch, duration, reverb, and / or rhythm of the target audio signal, or, if the target audio signal is binaural, changing the apparent sound source of the target audio signal according to the user, among one or more of these. In this last example, as the intermediate position changes, the apparent sound source of the target audio signal may be "panned" relative to the user. In such a case, the user may be asked to identify the target position within the training environment where the apparent sound source of the target audio signal appears to be emitted from a sound source in front of the user. For example, the apparent sound source of the target audio signal may initially be located to the right or left of the user, behind the user, and / or far away from the user. Based on intermediate input, the apparent sound source may be moved relative to the user, and the user may also be asked to identify where the apparent sound source of the target audio signal is located in front of and / or near themselves. Changing the target audio signal may be performed by the user device (e.g., a processor included within the user device) or by an external device or system (e.g., a remote or cloud-based system communicating with the user device).

[0030] The target location may be static or may change within the training environment. In particular, periodically or continuously changing the target location over time can increase the difficulty of the hearing training.

[0031] In a more preferred example, the information defined by the target audio signal may correspond to a target visual component within multiple different visual components in the training environment. Thus, the user may be asked to "identify" or select the appropriate visual component from within the training environment in response to listening to the target audio. Therefore, the method may include receiving user input corresponding to a visual component that the user believes to be related to the content of the target audio signal.

[0032] In a preferred example, the visual components may be buttons or displayed objects within the training environment. For example, the training environment may include a menu in a cafe, bar, or restaurant, and multiple visual components corresponding to the menu items offered by the cafe, bar, or restaurant. In response to hearing a target audio signal containing a customer's order, the user may be prompted to use a user interface to select one or more menu items desired by the customer. Alternatively, the training environment may include multiple animals forming a visual component, while the target audio signal may include animal sounds. In response to hearing animal sounds as part of the target audio signal, the user may be required to select the appropriate animal via a user interface.

[0033] Therefore, in such a training method, the user's task is to identify the correct visual component based on the content of the target audio signal. This identification task requires the user to distinguish the target audio signal from the background audio signal and select the appropriate visual component corresponding to the information within the target audio signal. This method helps train not only the user's working memory and attention but also their noise skills and clarity.

[0034] These types of identification tasks are designed to help users improve the clarity of their noise-hearing skills. Cognitively, users can develop selective attention and concentration, improving their ability to focus on specific objects or sounds. In particular, users may notice that this type of method improves social interaction, especially in crowded or noisy environments.

[0035] In a further example, the training method may include the user "matching" two distinct target audio signals. The user may be asked to evaluate whether the two target audio signals are similar and / or conceptually related. In a preferred example, the method may include sequentially providing the user with a target audio signal and a second target audio signal using an audio output, and receiving user input corresponding to the user's evaluation of whether the target audio signal and the second target audio signal are similar and / or related. Such a method can develop not only the user's memory but also their ability to distinguish sounds from one another.

[0036] In such examples, the method may include providing two or more target audio signals using an audio output, and user input indicating whether the user believes the two or more target audio signals are similar and / or related. In a particularly preferred example, each of the two or more target audio signals may be provided in response to the receipt of preliminary user input corresponding to that target audio signal. For example, in a "tile-matching" training method, the user may provide preliminary input regarding tiles or other selectable visual components, and accordingly, the user may be provided with corresponding target audio signals. If multiple tiles share similar (e.g., the same) target audio signals and / or related target audio signals, the user may provide user input indicating that they understand the corresponding tiles match.

[0037] These types of matching tasks are designed to help users develop their discriminative hearing skills. Cognitively, they can improve users' short-term or working memory, including both auditory working memory and visuospatial short-term memory (particularly in the "tile-matching" example mentioned above). Thus, this type of approach is designed to benefit users' reading comprehension, concentration, and language learning abilities.

[0038] Therefore, as described above, the method can include a variety of tasks such as foraging / seeking, content identification, and matching. In a preferred example, these different tasks form alternative modes of the method for performing hearing training. For example, the method may include performing one or more of the foraging / seeking mode, the identification mode, and / or the matching mode, each of which may take the form described above. In such an example, the method may include performing one or more of these training modes based on user input and / or the results of a standardized hearing test performed by the user in a preliminary method step. Therefore, it will be understood that as the mode of hearing training changes, the type or format of information defined by the task required of the user and / or the target audio signal changes.

[0039] Standardized hearing tests may include: AIADH (Amsterdam Inventory for Auditory Disability and Handicap) ("Subjective Hearing Impairment Factors," Kramer, Kapteyn, Festen, and Tobi, Audiology, November-December 1995; 34(6):311-20) (a series of multiple-choice questions in which a user assesses how hearing affects their quality of life); HearWHO (World Health Organization, 2018) (a test in which a user is asked to hear and identify digits spoken in a voice test with background white noise or other noise); a test of the highest frequency a listener can hear; a pure-tone audiometry test that tests the threshold volume at which a listener can hear sounds across a range of frequencies (such as the approach defined in ISO 8253-1:2010 issued by ISO 2010-11); or one or more of other appropriate tests that test a user's hearing ability.

[0040] In a preferred embodiment, the method further includes the step of analyzing user input received at a user interface to determine whether the user evaluation indicated by the user input corresponds to information defined by the target audio signal. Thus, the method may include determining whether the user has correctly identified the information conveyed by the target audio signal. Preferably, the feedback provided to the user is based on the results of this analysis step. This analysis may be performed by the user device (e.g., by a processor within the user device) or by a further device or system outside the user device (e.g., a remote or cloud-based system communicating with the user device).

[0041] Preferably, the method includes iteratively repeating the method steps described herein. Thus, the user can repeatedly perform the training to improve their listening skills. Thus, the user can perform listening training multiple times within a training session, and the target audio signal (and preferably the information to be determined) and / or background audio signal are changed between each iteration of the method steps. Thus, in a preferred example, the method includes iteratively repeating the steps of the method for a single user, i.e., ensuring that user input received in different iterations is received from the same user.

[0042] In a preferred example, the method may include repeatedly performing the steps of the method for a predetermined period (e.g., a period of 3 to 30 minutes), a predetermined number of repetitions (e.g., 10, 20, 30, or 50 repetitions), or until the user can no longer provide user input that precisely corresponds to the information defined by the target audio signal. If the method is repeated a predetermined number of times, the predetermined number may be in the range of 5 to 50, and more preferably in the range of 10 to 30.

[0043] In such examples, the method may further include providing the user with session feedback based on the user's performance over multiple iterations of the method for hearing training. For example, session feedback may include instructions for the total number of iterations in which user input provided by the user was received that correctly corresponds to information defined in the target audio signal, the percentage of iterations in which correct user input was received, and / or the highest number of consecutive iterations in which the device received correct user input (e.g., the highest number of consecutive correct answers provided by the user). Session feedback may include any of the following: audio instructions, visual instructions, and / or haptic instructions (e.g., vibration of the user device), or other features of the feedback provided after each iteration as described above. Session feedback may be generated by the user device (e.g., a processor within the user device) and / or any other device or system (e.g., a remote or cloud-based system communicating with the user device).

[0044] More preferably, the difficulty level of the listening training may be increased in subsequent iterations of the method based on the determination that the user evaluation indicated by the user input corresponds to the information defined by the target audio signal. Thus, this step may be based on the results of the analysis process described above. In this way, the difficulty level of the listening training can be increased as the user correctly identifies the information conveyed by the target audio signal. Thus, the user's listening skills may be further developed through more difficult training as their listening improves. Furthermore, or alternatively, the difficulty level of the listening training may be decreased in subsequent iterations of the method based on the determination that the user evaluation indicated by the user input does not correspond to the information defined by the target audio signal. By changing the difficulty level of the listening training, the training can be customized to the user and individual training outcomes can be improved.

[0045] In other words, a single user's performance across different iterations of listening training can be used to adjust the difficulty of future training sessions. Adaptive difficulty maintains user engagement and challenges and helps users continue to grow their listening skills over time.

[0046] The difficulty of subsequent iterations of the method may vary depending on the user's success and / or failure. Alternatively, the difficulty may be periodically changed based on whether the user has met a threshold for success or failure in user input over a series of consecutive iterations of the method (e.g., a predetermined number of consecutive successes and / or failures, or a predetermined percentage of successes and / or failures over a series of consecutive iterations of the method). "Success" is understood to mean an iteration in which the user input received from the user correctly corresponds to the information defined by the target audio signal (i.e., the user correctly evaluated or identified the information in the target audio signal). "Failure," on the other hand, means that the user input does not correspond to the information defined by the target audio signal. Therefore, the difficulty can be automatically adjusted to suit the user's needs.

[0047] The method involves determining the proportion of user evaluations represented by user inputs that accurately correspond to the information defined by each target audio signal over multiple consecutive iterations of the method. If the proportion of correct user inputs is greater than a predetermined first value, the difficulty of the listening training increases in one or more subsequent iterations; or if the proportion of correct user inputs is less than a predetermined second value, the difficulty of the listening training decreases in one or more subsequent iterations. Therefore, the difficulty may be automatically adapted to reflect user performance by comparing the user's success and failure rates with a predetermined threshold.

[0048] In particular, the inventors recognized that user engagement increases significantly when the user correctly evaluates the information in the target audio signal in about 85% of iterations (i.e., a success rate of about 85% and a failure rate of about 15%). If the user is wrong frequently, they may become frustrated, and if the user is right very frequently, they may find the training boring.

[0049] Therefore, a predetermined first value beyond which training becomes more difficult may be 95% or higher, and more preferably 90% or higher. Similarly, a predetermined second value beyond which the difficulty of training decreases may be 70% or lower, and more preferably 80% or lower. Therefore, the success rate without changing the difficulty is preferably 70-95%, and more preferably 80-90%.

[0050] In a preferred example, hearing training may consist of a series of training sessions or rounds, each containing multiple consecutive repetitions of the method. For example, each training session may contain 5 to 50 repetitions, more preferably 10 to 30 repetitions. The percentage of user ratings that correctly relate to each piece of information in the sequence of target audio signals throughout the training sessions is determined at the end of each of these training sessions, and the difficulty of subsequent training sessions (i.e., rounds) may be adjusted based on this determination. As mentioned above, in some examples, the number of repetitions for each training session may be predetermined, but this is not required. For example, the number of repetitions for a training session may be defined by the number of repetitions that the user can complete within a time limit.

[0051] In a particularly preferred example, the method may include a preliminary step of setting a baseline difficulty based on the aggregated performance of the user across previous training sessions, which include multiple iterations. This may be done using predetermined first and second values, as described above. During subsequent training sessions, the difficulty may be adjusted or adapted from this baseline difficulty based on the results within the training session. Thus, the difficulty of each iteration of training will depend on both the user's performance in previous training sessions and the user's performance in previous iterations of the method in the ongoing (i.e., concurrent or current) training session.

[0052] Alternatively, after each iteration of the method, the percentage of user ratings indicated by user inputs that precisely correspond to the information defined by each target audio signal may be calculated for the group prior to the iteration (e.g., for the previous 20 or 30 iterations). Thus, the success rate of the most recent iteration of the method is calculated iteratively, and the difficulty is adjusted in a rolling manner. For example, the group of iterations on which the difficulty is based may include at least 10 previous iterations, more preferably at least 15, 20, or 30 previous iterations.

[0053] In addition, or instead, the method may include a preliminary step of conducting a standardized hearing test, and based on the results of the standardized hearing test, the difficulty of the hearing training may be changed, the content and / or one or more characteristics of the target audio signal and / or background audio signal may be changed, and / or the mode of the hearing training may be changed. For example, the frequency and / or volume of the target audio signal and / or background audio signal may be changed depending on the user's results in the standardized hearing training. For example, the frequency and volume of the target audio signal and / or background audio signal may be changed based on the results of a standardized test performed by the user. This helps ensure that the user can hear the audio signal and that their hearing skills are effectively trained. Similarly, the mode of the hearing training performed may be changed depending on the user's performance in the standardized hearing test, such as the user being required to perform different tasks and / or the format in which the target audio signal provides information to the user being different. Thus, hearing training may be customized to a particular person using the training method described herein.

[0054] For example, standard hearing tests may include (as mentioned above) AIADH (Amsterdam Inventory for Auditory Disability and Hnadicap), HearWHO (World Health Organization, 2018), tests for the highest frequency a listener can hear, pure-tone hearing tests that test the threshold volume at which a listener can hear sounds across various frequencies (such as the approach defined in ISO 8253-1:2010 issued by ISO 2010-11), or other appropriate tests to test a user's hearing skills. In any case, training can be made more effective by being customized to the user.

[0055] There are various approaches to automatically adjusting the difficulty of listening training. For example, increasing the difficulty of listening training can be done by decreasing the volume of the target audio signal relative to the background audio signal, degrading the quality of the target audio signal compared to the background audio signal (e.g., by applying a bandpass, lowpass, or highpass filter to the target audio signal), increasing the similarity between the target audio signal and the background audio signal (e.g., by providing sounds in the background audio signal that are of similar frequency to the target audio signal or that appear to the user to originate from similar sound sources), increasing the number of sounds in the background audio signal, and, if the target audio signal includes binaural audio, changing the apparent position of sound sources in the target audio signal relative to the user during repetitions, and / or the successive repetitions of the method. This may include one or more of the following: increasing the variation in the position of one or more sounds within the target audio between iterations; increasing the variation in the apparent position of one or more sound sources for each sound in the background audio relative to the user during iterations, if the background audio signal includes binaural audio; increasing the complexity of the target audio signal and / or background audio signal (for example, by using sounds that are more difficult for the user to distinguish or identify, by shortening the duration of the target audio signal, by increasing the speaking speed of the voices in the target audio signal, by using sounds that contain multiple pieces of information in the target sound that the user should identify with the same or multiple user inputs); imposing time limits on when the user must provide user input; and / or increasing the visual complexity of the training environment displayed. Most of these examples make it more difficult for the user to distinguish the target audio signal from the background audio signal and / or make it more difficult for the user to identify the information defined by the target audio signal.In fact, examples of changing the relative volume and audio quality of the target audio signal and background audio, or increasing the signal level, can be understood as equivalent to changing the SNR (Signal to Noise Ratio) provided to the user. On the other hand, increasing the visual complexity of the displayed training environment can easily distract the user, requiring them to concentrate and pay more attention during training. It should also be understood that the difficulty of listening training can be reduced by the opposite approach, by performing the opposite of one or more of the above options.

[0056] Preferably, the difficulty level may be adjusted incrementally after each iteration and / or training session. This may reflect gradual or incremental improvements in user feedback as the user continues with the method. Each of these incremental steps of increasing or decreasing the difficulty level may include one of the actions described above. In a particularly preferred example, the changes are prepared so that the difficulty level increases incrementally.

[0057] In some embodiments of the method, the difficulty of the training, and therefore the user's ability to accurately identify information in the target audio signal from the content of the background audio signal, may be quantified using the signal-to-noise ratio. Thus, the signal-to-noise ratio may be presented to the user or expert as a score quantifying performance in hearing training, allowing for tracking of performance over time.

[0058] The examples of how the difficulty level of the hearing training described above can be varied are common to all embodiments described above. However, the difficulty level of each of the different potential modes of the present invention described above can be modified in more specific ways.

[0059] For example, the "foraging" / "seeking" methods described above may reduce the number of properties that would otherwise be difficult to modify the target audio based on the relative position of the intermediate and target locations. Consequently, less information about the location of the target location is provided to the user. In addition, or instead, the location of the target location may be moved as described above. In addition, or instead, the user may be required to provide more precise user input (i.e., the user input must be close to the target location) in order to be considered to correctly correspond to the target location.

[0060] The "identification" method described above may increase the difficulty by providing multiple content items within each target audio signal, requiring the user to correctly identify each item using user input. For example, if the training environment is a cafe, bar, or restaurant, the target audio signal might be "Can I have a black coffee and a croissant, please?", and the user may be required to provide input related to both black coffee and a croissant (i.e., select the visual components corresponding to both products). Similarly, different target audio signals and visual components used within the method may be made more similar. For example, it may be more difficult for a user to distinguish between "carrot cake" and "caraway cake" than between "carrot cake" and "lemon cake". Likewise, customers may place orders from various positions relative to the user; for example, the apparent sound source of the target audio signal corresponding to the customer's voice may pan from left to right or up and down relative to the user. The customer may speak faster or make more ambiguous requests. As the difficulty increases, additional distracting sounds within the background target signal may include one or more other customers waiting in line, passing traffic (cars, buses, trucks, etc.), or other sounds from within the cafe. The apparent location of these background sounds may also vary within each iteration of the listening training or between different iterations. A bandpass filter (for example, a filter configured to reduce the volume of a specific frequency up to 5kHz or 3kHz) may be applied to the target audio signal to simulate a face mask worn by a customer.

[0061] The matching method described above may increase the difficulty by making the non-matching target audio signals more similar. For example, if the user needs to determine which target audio signals are identical, the content of the different target audio signals may be made more similar (e.g., including rhyming words or words that differ in fewer letters or syllables), or they may include tones that are similar in pitch, reverb, duration, and / or rhythm.

[0062] In a further example, the difficulty level of the training method may be modified by the user, i.e., the difficulty level may be changed in response to input from the user. This modification of difficulty level may include any of the modifications described above. These modifications of difficulty level may occur between consecutive iterations of the method steps or during a single iteration of the method. For example, if a user is having difficulty distinguishing between a particular target audio signal and a background audio signal during an iteration of the method, the user may provide input requesting that the target audio signal be repeated, that the background audio signal be removed or reduced in volume compared to the target audio signal, that the target audio signal be displayed, and / or that the audio within the target audio signal be displayed (e.g., using subtitles). Allowing users to manipulate the training method in this way makes it easier to customize listening training to the user, reduces user frustration, and improves user engagement with listening training.

[0063] Preferably, the user device includes a display, and the method includes using the display to show the training environment to the user. For example, the training environment may include images, videos, augmented reality, and / or virtual reality. Showing the training environment to the user further increases the sensory input to the user during hearing training. This enhances the sense of reality of the hearing training, as users typically experience both visual and auditory input in their daily lives. Thus, the effectiveness of the hearing training is enhanced. However, this is not mandatory, and in further embodiments, the hearing training may include providing the user with only audio signals.

[0064] Displaying a training environment may increase the complexity of hearing training. Users may be asked to associate information defined by the target audio signal with visual components displayed within the training environment. For example, users may be asked to select a location or item displayed within the training environment.

[0065] In a similarly preferred embodiment, the training environment includes VR (Virtual Reality) and / or AR (Augmented Reality), and the method includes changing the apparent sound sources in binaural audio of a target audio signal and / or background audio signal in response to a change in the viewpoint of the training environment displayed by the user device.

[0066] If the training environment includes VR (Virtual Reality), the viewpoint from which the training environment is displayed may change as the user moves their head. The virtual reality system tracks where the user is looking and adjusts the viewpoint displayed to the user through the virtual reality headset accordingly. Similarly, if the training environment includes augmented reality, the viewpoint from which the training environment is viewed may change as the device displaying the augmented reality training environment (e.g., smartphone, tablet, or headset) moves. Therefore, the device performs binaural synthesis to achieve accurate spatialization of sound in real time according to the user's viewpoint.

[0067] Binaural audio within target and / or background audio signals may be modified based on changes in viewpoint due to the use of a Head-Related Transfer Function (HRTF), which changes according to the angle between the assumed user's head and the apparent sound source. Therefore, the apparent sound source within the binaural target and / or background audio signals can be accurately spatialized throughout the entire use of the AR / VR training environment. As a result, the spatialized audio can be maintained even when the viewpoint of the training environment displayed to the user changes. Spatialized audio is particularly realistic and accurately reflects the user's experience beyond VR / AR hearing training.

[0068] Preferably, the target audio signal includes one or more of the following: human voice, animal sounds, traffic noise, musical instruments, nature sounds, ambient noise, or synthesized sound effects. However, any suitable sound effect or recorded sound can be used in the target audio signal. As mentioned above, the target audio signal preferably includes binaural audio, so that the sounds described above may have apparent positions relative to the user depending on the different signals presented to the user's different ears.

[0069] Preferably, the background audio includes one or more of the following: human voices, animal sounds, traffic noise, musical instruments, weather noise, water sounds, nature sounds, synthesized sounds, ambient noise, white noise, or synthesized sound effects. In further embodiments, any suitable sound effect or recorded sound can be used within the background audio signal.

[0070] More preferably, background audio includes multiple sounds that overlap at least partially. This layering of sounds provides the user with multiple distinct sounds simultaneously, creating a realistic soundscape. Particularly valuable examples of background audio signals include both longer, relatively quiet ambient sounds and louder, shorter, distracting sounds. For example, multiple human conversations can be combined to create the noise of a group in a cafe, bar, or restaurant. A rainforest, on the other hand, may be simulated by layering animal calls with sounds such as dripping water and rustling leaves in the wind. For the user, distinguishing the target audio signal from a complex soundscape containing multiple overlapping sounds, and then identifying the information defined or conveyed by the target audio signal, is a particularly challenging test. This can improve the effectiveness of listening training. Indeed, as mentioned above, if the background audio signal preferably includes binaural audio, each of the overlapping sounds may have a different apparent position relative to the user. The arrangement of multiple background sounds around the user is particularly realistic and helps improve training outcomes.

[0071] Using overlapping sounds or ambient recordings as background audio signals (especially binaural sounds and recordings) provides a highly realistic soundscape and yields improved training results compared to using random noise signals such as white noise, pink noise, and brown noise. White noise is a random signal that has equal intensity at different frequencies and gives a constant power spectral density. Pink noise, or 1 / f noise, is a random noise signal with a power spectral density inversely proportional to the signal frequency. Brown noise (also called red noise) is a random noise signal with a power spectral density inversely proportional to the square of the signal frequency. While white, pink, and brown noise are constant and easy to generate, they do not reflect natural environments and are unrealistic substitutes for ambient sounds that users experience in everyday life.

[0072] Preferably, the target audio signal and / or background audio signal are within the range of human hearing. For example, each of the target audio signal and / or background audio signal may include sounds between 20 and 20,000 Hz, more preferably between 25 Hz and 15,000 Hz, and even more preferably between 100 and 10,000 Hz. The volume of the target audio signal may be varied relative to the background audio signal (for example, to change the difficulty of training), but preferably the background audio signal is quieter than the target audio signal so that the user can identify the target audio and determine the information conveyed by the target audio signal. For example, the volume of the background audio signal may be at least -6 dB, more preferably at least -12 dB, relative to the target audio signal.

[0073] Preferably, the user device is a smart user device, and preferably, the user device is a smartphone, tablet, laptop, personal computer, or AR and / or VR system. In contrast to conventional systems, which are often installed in laboratories, this type of personal device is easily accessible to the user. Smartphones and tablets are particularly portable and convenient for the user. On the other hand, using an AR and / or VR system, including a headset, can provide more complex training settings. Not limited to these, any suitable user device can be used.

[0074] Preferably, the user device includes a pointing device, which is a user interface that allows the user to provide spatial data to the user device. For example, the user device may include a touchscreen, trackpad, mouse, mousepad, joystick, or gamepad. However, this is not required, and any suitable input device may be used. For example, the input device may be a microphone, and the user may provide audio input (e.g., voice input).

[0075] In a particularly preferred example, the user device's display and user interface may be combined. For instance, the user device may have a touchscreen. This is a particularly space-efficient and intuitive means for the user to interact with the user device.

[0076] Preferably, the target audio signal and background audio signal are provided to the user via headphones or an alternative audio output device connected to the audio output. The term headphones will be understood to include earphones, earbuds, headsets, and any other suitable form of audio output device worn on the user's head. Headphones provide a particularly convenient means of providing separate left and right audio channels of binaural audio directly to the user's corresponding ears. In alternative embodiments, the method may also include providing the target audio signal and background audio to the user via a system of loudspeakers configured to provide separate left and right audio channels of binaural audio to the user's corresponding ears.

[0077] A further aspect of the present invention provides a computer implementation method for performing hearing training using a user device comprising a user interface and an audio output, the method comprising: providing a target audio signal using the audio output that defines information to be determined by the user; the target audio signal including binaural audio; receiving user input at the user interface corresponding to the user determination of the information defined by the target audio signal; and providing feedback to the user based on the result of the user determination.

[0078] Such methods also provide a realistic and effective way to train users' hearing. The use of binaural audio within a target audio signal mimics sounds and situations that users experience in everyday life.

[0079] Methods according to this aspect of the present invention have any of the features described above with reference to a previous aspect of the present invention and can provide corresponding advantages including any preferred features described above. For example, according to this aspect of the present invention, a background audio signal is not required, but in a preferred embodiment, a background audio signal is provided. This improves the realism and effectiveness of hearing training because the user needs to distinguish the target audio signal from the background audio before interpreting the information provided by the target audio signal. The background audio signal may include binaural audio or conventional mono and / or stereo audio.

[0080] According to a further aspect of the present invention, a user device is provided comprising a user interface and an audio output, the user device being configured to perform a hearing training method according to any of the above-described aspects of the present invention.

[0081] The user device may comprise any of the physical components described above with reference to the preceding aspects of the present invention and may be configured to perform any of the preferred or optional method steps described above. Such a user device provides the advantages corresponding to the examples described above.

[0082] According to a further aspect of the present invention, a non-temporary computer-readable medium is provided that, when read by a processor, stores instructions causing a user device to execute a hearing training method according to any of the methods described above.

[0083] When the instruction is read by the processor, it may cause the user device to perform any of the preferred or optional method steps described above. Such an instruction provides the advantages corresponding to the examples described above.

[0084] A specific example of the present invention will be described with reference to the following diagram. [Brief explanation of the drawing]

[0085] [Figure 1] A schematic diagram showing a system equipped with a user device according to the present invention. [Figure 2] A flowchart illustrating the method according to the present invention. [Figure 3a] A schematic diagram showing a user device for performing the method according to the present invention. [Figure 3b] A schematic diagram showing a user device for performing the method according to the present invention. [Figure 3c] A schematic diagram showing a user device for performing the method according to the present invention. [Figure 4] A schematic diagram showing a user device for performing the method according to the present invention. [Figure 5] A schematic diagram showing a user device for performing the method according to the present invention. [Figure 6] A flowchart illustrating the method according to the present invention. [Modes for carrying out the invention]

[0086] Figure 1 schematically shows a system 1 comprising a user device 10 configured to perform a method for hearing training. The user device 10 may be a smart user device such as a smartphone, tablet, laptop, or personal computer. The user device 10 comprises a processor 11, memory 12 (i.e., computer-readable storage medium), a user interface 13, a display 14, and an audio output 15. In practice, the user device may have further features not shown in this schematic diagram.

[0087] The processor 11 is configured to execute instructions stored in the memory 12 of the user device 10. The user interface 13 is configured to receive input from the user (i.e., to receive user input). The display 14 is configured to display the training environment to the user. The user device 10 may have a touchscreen that provides both the user interface 13 and the display 14. Alternatively, the user interface 13 and the display 14 may be separate components. For example, the user interface 13 may include a touchscreen, trackpad, mouse, mousepad, joystick, gamepad, or any other suitable input device.

[0088] The user device 10 is configured to connect to an external audio output device, such as headphones 21 or a loudspeaker 22, via an audio output 15. The connections 21a, 22a between the user device 10 and the headphones 21 and / or loudspeaker 22 may be wired or wireless (e.g., via Bluetooth®, Wi-Fi®, or any other suitable alternative wireless communication protocol). The user device 10 may use the audio output 15 and these connections 21a, 22a to provide audio signals, which are then converted into audio (i.e., voice) by the headphones 21 or loudspeaker 22.

[0089] In particular, the user device 10 is configured to provide binaural audio to the user through headphones 21 or speakers 22. Binaural audio consists of left and right audio channels, and the difference between the left and right audio channels is based on the assumed relationship between the listener's ears (for example, defined by HRTF (Head Related Transfer Function)). The user device 10 may provide binaural audio that is binaurally recorded (i.e., recorded using a pair of microphones placed on either the head of a model or the head of a person) or generated from sampled signals using a head-related transfer function.

[0090] The user device 10 shown in Figure 1 is suitable for use in the manner shown by the flowchart in Figure 2.

[0091] In step s101, the user device 10 provides at least a target audio signal using its audio output. Preferably, the user device 10 also provides a background audio signal that at least partially overlaps with the target audio signal (for example, so that at least a portion of the sound in the target audio and the background audio are provided simultaneously). The target audio signal defines information to be determined by the user. At least one of the target audio signal and the background audio signal includes binaural audio.

[0092] During training, the target audio signal and any background audio signal are provided to the user by the audio output 15 of the user device 10 (e.g., via headphones 21 or loudspeaker 22). Binaural audio may be binaurally recorded and / or generated from sampled signals using a head-related transfer function. Optionally, during this step, the binaural audio can be adapted or dependent on the relative position and orientation between the user's head and the apparent sound sources in the binaural audio. The apparent sound sources of one or more sounds in the target audio signal and / or background audio signal may change depending on the position and orientation of the user's head or user device relative to the apparent sound sources. To achieve this, the position and orientation of the user's head may be tracked.

[0093] When listening to the target audio, the user attempts to identify the information defined by the target audio. If background audio is present, the user must first distinguish the target audio from the background audio. Subsequently, the user provides user input to the user device 10 using the user interface 13, corresponding to the understanding or evaluation of the information conveyed or given by the target audio. Thus, in step s102, the user device 10 receives user input corresponding to the user's evaluation of the information defined by the target audio signal.

[0094] Upon receiving user input (s102), the user device 10 provides feedback to the user in step s103 based on the user evaluation indicated by the user input. Thus, the user receives an evaluation of their ability to understand and interpret the audio provided by the user device 10. Consequently, the user can train and improve their hearing.

[0095] In order to provide feedback in step s103, user input received by the user interface may be analyzed to determine whether the user evaluation corresponds to information defined by the target audio signal. This analysis may be performed by the processor 11 of the user device 10 or any other suitable processor. If the user evaluation indicated by the user input accurately corresponds to information defined by the target audio signal (i.e., the user correctly identifies the information conveyed by the target audio), the user device 10 may receive positive feedback. Otherwise, the user device 10 may provide negative feedback. The feedback may be in the form of a visual indicator, such as a message displayed on the display 14 of the user device 10; an audio indicator, such as a sound effect or verbal message provided by the audio output 15 of the user device 10; and / or any other suitable indicator, such as a tactile indicator, such as vibration that can be generated using a vibration unit in the user device 10.

[0096] In a preferred example, method steps s101, s102, and s103 are repeated iteratively to allow the user to continue training and developing their listening skills. The difficulty level of the listening training may be adjusted incrementally based on the user's (i.e., a single user's) success and / or failure in previous iterations of the method. In addition, or alternatively, the difficulty level may be adjusted according to user input. In addition, or alternatively, the difficulty level of the listening training may be based on the results of a preliminary step in which a standardized listening test (e.g., Amsterdam Inventory, HearWHO, or a test of the highest frequency the user can hear) is administered to the user. The user device 10 may be configured to administer such a standardized listening test via audio output 15. However, in other examples, the results of the standardized listening test may be received by the user device 10 from an external device or system. Examples of how the difficulty levels of different training methods can be changed or manipulated are described in the summary section above.

[0097] Specific examples of methods for conducting hearing training using user devices 30, 40, and 50 are described with reference to schematic Figures 3 to 5. Each of these examples incorporates the steps described above with reference to Figure 2.

[0098] User devices 30, 40, and 50 are smartphones equipped with touchscreens 31, 41, and 51 that provide both a display and a user interface. User devices 30, 40, and 50 are equipped with audio outputs (not shown) configured to provide a target audio signal and / or a background audio signal to the user (e.g., via headphones or a loudspeaker array). In either case, one or both of the target audio signal and the background audio signal may constitute binaural audio. In addition, user devices 30, 40, and 50 may share any of the further features of user device 10 described above with reference to Figures 1 and 2. Interactions between the user and user devices 30, 40, and 50 are shown by hand icons in Figures 3a–3c and 4, and by hatched visual components in Figure 5.

[0099] Figures 3a and 3b schematically illustrate a series of steps in a foraging or seeking hearing training method performed using the user device 30. The user device 30 displays the training environment 32 to the user using a touchscreen 31. Within the training environment 32, a hidden target location 33 is defined that is unknown to the user at the start of the hearing training.

[0100] Figure 3a shows how the target audio signal provided by the user device 30 can be modified based on the distance d between the intermediate position L and the target position 33, using the audio output of the user device 30. On the other hand, Figure 3b shows the movement m of the user input through the training environment 32 as the user searches for the target position 33.

[0101] The user device 30 provides the user with a target audio signal and, preferably, background audio throughout the method. The audio signals provided to the user depend on preliminary user input received from the user via the touchscreen 31, which corresponds to intermediate positions L in the training environment (indicated by hand icons in Figures 3a-3c). In particular, one or more characteristics of the target audio signal change depending on the position of intermediate position L relative to the hidden target position 33. Thus, the target audio signal conveys information about the position of the target position 33 in the training environment 32.

[0102] Specifically, as shown in Figure 3a, the user device 30 receives a series of preliminary user inputs on the touchscreen 31 that correspond to a series of intermediate positions L1, L2, L3, and L4 within the training environment 32. For example, the user may tap or drag their finger on the touchscreen 31 at each of the intermediate positions L1, L2, L3, and L4. In other words, as the user attempts to locate the target position 33, the intermediate positions L1, L2, L3, and L4 provided by each preliminary user input change as indicated by the arrows m1, m2, m3, and m4.

[0103] In response to each preliminary user input, the user device 30 calculates the distances d1, d2, d3, and d4 between intermediate positions L1, L2, L3, and L4 and the target position 33, and modifies the characteristics of the target audio signal accordingly. For example, the volume of the target audio signal may be changed (for example, the target audio may become louder or quieter as it approaches the target position 33). In addition, or instead, the content, pitch, duration, reverb, or rhythm of the target audio may be changed. For example, if intermediate position L is close to the target position 33, the pitch or tempo of the target audio signal may be increased. In addition, or instead, if the target audio signal is binaural, the apparent sound source of the target audio signal may change to the user.

[0104] As shown in Figure 3b, after inputting a series of preliminary user inputs corresponding to intermediate positions L1, L2, L3, and L4, and listening to the resulting changes in the target audio, the user can evaluate or determine where the target position 33 is located. The user then determines the position L of the target position 33 in the training environment 31. * User input corresponding to the evaluation may be provided (for example, by double-tapping the touchscreen 31 of the user device 30).

[0105] Subsequently, the user device 30 moves to the position L indicated by the user input. * The system may determine or analyze whether the input corresponds to the target position 33 and provide feedback to the user based on the results of this analysis. As shown in Figure 3b, the position L indicated by user input * The input is accurate to the target location 33 (for example, within a predetermined distance from the target location), and therefore the user can provide positive feedback. However, if the user inputs a location far from the target location 33, the user may receive negative feedback. This feedback helps improve the user's listening skills. The feedback may also be based on the time or number of intermediate locations required for the user to identify the target location.

[0106] Specific target audio signals suitable for use in this "foraging" / "seeking" method include animal calls and sounds (such as bird calls) that can be heard against a forest background audio signal soundscape, including separate overlapping sounds such as wind rustling through leaves, flowing water, and animal calls. Similarly, the target audio may be the sound of a frying pan (or other cooking utensil) heard against background audio or the sounds of a noisy kitchen or market.

[0107] Figures 3a and 3b will show a series of discrete intermediate positions L1, L2, L3, and L4, indicated by a series of corresponding discrete user inputs. However, this is not mandatory, and the user may input a continuous range of intermediate positions (for example, by dragging a finger on the touchscreen 31). In this case, the target audio signal provided to the user may change continuously.

[0108] Furthermore, while the position of the target position 33 is static in Figures 3a and 3b, in further examples, the position of the hidden target position 33 may change periodically or continuously.

[0109] As described above in relation to Figures 3a and 3b, one or more characteristics of the target audio signal may vary depending on the magnitude of the distance between each intermediate position L1, L2, L3, L4 and the target position 33, but this is not mandatory. Instead, as shown in Figure 3c, the characteristics of the target audio signal may vary depending on the vertical distance v1, horizontal distance h1, and / or angle θ1 between the intermediate position L1 and the target position 33, as indicated by the user input.

[0110] In some examples, different characteristics of the target audio may be varied based on each of these different coordinates, which can quantify the relative position between the intermediate position L and the target position 33. For example, the volume of the target audio signal relative to the background audio signal may vary depending on the vertical distance between the intermediate position L and the target position 33, while the apparent sound source of the binaural audio within the target audio signal may change relative to the user depending on the horizontal distance between the intermediate position L and the target position 33. In this example, the user may need to identify where the volume of the target audio is loudest and where it appears to be coming directly from a sound source in front of them.

[0111] Figure 4 schematically illustrates a method for hearing training performed using a user device 40 (smartphone), in which the user must "identify" the content of the target audio signal.

[0112] As shown in Figures 3a to 3c, the user device 40 displays the training environment 42 to the user using a touchscreen 40. The hearing training shown in Figure 4 includes identifying orders placed by customers in a cafe, bar, or restaurant. The training environment 42 displayed by the user device 40 is divided into two sections: a customer section 42a where a customer making an order may be displayed, and a menu section 42b where multiple selectable visual components 43 are displayed corresponding to different products on the cafe, bar, or restaurant menu.

[0113] When customer C is displayed by user device 40, user device 40 provides the user with a target audio signal corresponding to customer C's order (e.g., via headphones). For example, the target audio signal may include speech such as "I'd like a black coffee" or "Could I have a slice of apple cake?". Thus, the user needs to identify the product that customer C desires and select the corresponding visual component 43. For example, the user may provide user input corresponding to the visual component by tapping the touchscreen of user device 40, as indicated by a hand icon 44. Therefore, the information defined by the target audio signal is the linguistic content of human speech within the target audio, while the user input is the selection of a visual component 43 that the user believes corresponds to this content of the target audio signal.

[0114] Upon receiving user input, the user device 40 may provide feedback to the user based on the user input in order to improve the user's listening skills. Prior to this, the user device 40 (or another device or system) may determine whether the visual component 43 indicated by the user input accurately corresponds to the content of the target audio signal. The feedback may also be based on the rate at which the user provides user input.

[0115] The user device 40 may provide background audio, such as ambient noise in a bar, cafe, or restaurant, which may be provided simultaneously with the target audio signal, i.e., the background audio signal and the target audio signal may overlap. As mentioned above, a particularly realistic soundscape can be created if the background audio signal includes multiple overlapping sounds, such as conversations between multiple people, the noise of a coffee machine, cutlery and dishes, and / or traffic noise. At least one of the background audio signal and the target audio signal includes binaural audio.

[0116] In the example described above, the training environment displays customers in a cafe, bar, or restaurant, and the content of the target audio signal that the user needs to identify is the linguistic content of human speech (i.e., the actual words spoken by the customers). However, this is not mandatory, and in other examples, the target audio signal and training environment may take other forms. For example, the training environment displays a farm, and the target audio signal includes the sounds of farm animals. In this case, the user may be required to identify the appropriate livestock displayed by the user device 40 from their sounds. In such an example, the background audio may include typical sounds heard on a farm.

[0117] Figure 5 schematically illustrates a method of hearing training performed using a user device 50 (smartphone), in which the user must "match" different target audio signals together.

[0118] The user device 50 uses its touchscreen 51 to display a training environment 52 that includes multiple visual components 53 that can be selected by the user. Specifically, as shown in Figure 5, the selectable visual components 53 take the form of tiles, and the user can select each tile by tapping the touchscreen 51 on it. In this way, the user provides the user device 50 with preliminary user input corresponding to the visual components 53.

[0119] For example, user device 50 may receive a first preliminary user input corresponding to a first visual component 53a (shown hatched in Figure 5) and provide the user with a first target audio signal corresponding to the first visual component 53a via its audio output (not shown). Subsequently, user device 50 may receive a second preliminary user input corresponding to a second visual component 53b (shown hatched in Figure 5) and provide the user with a second target audio signal corresponding to the second visual component 53b via its audio output (not shown). After listening to both target audio signals, the user needs to evaluate whether the first and second target audio signals are similar and / or related, that is, whether the target audio signals corresponding to the selected visual components 53a and 53b match. For example, matching target audio signals may be identical and / or share similar or the same audio characteristics such as pitch, rhythm, duration, timbre, and / or reverb. Alternatively, matching target audio signals may be conceptually linked; for example, the first target audio signal may contain human speech saying the term "dog," while the second target audio signal may contain a dog barking. Alternatively, or additionally, if the target audio signals are binaural, the user may need to determine whether the target audio signals share the same apparent sound source, i.e., whether the target audio signals are similarly spatialized to the user. In this way, the information provided to the user by the target audio signals is their relationship and / or similarity to other target audio signals.

[0120] If the user believes that target audio signals corresponding to two or more different visual components 53 are similar and / or related (for example, the first and second visual components 53a and 53b shown in Figure 5), they may provide user input corresponding to those different visual components 53. For example, the user may "double tap" each of the visual components 53 shown in the training environment and / or "drag" one of the visual components 53 onto the other visual component 53.

[0121] When the user device 50 receives user input from a visual component that the user believes to be linked by a corresponding target audio signal, it may provide feedback to the user based on this user input. For example, if the user correctly identifies a visual component that shares a relevant and / or similar target audio signal, the user may receive positive feedback.

[0122] In addition to the target audio signal, the user device 50 may provide the user with a background audio signal via an audio output (not shown). In these methods, the user needs to distinguish the target audio from the background audio before they can begin comparing different target audio signals. One or both of the target audio signal and the background audio signal may include binaural audio.

[0123] Following the explanation of Figure 5 above, it will be understood that various similar training methods may be used with matching techniques, and that the training methods are not limited to the "tile-matching" approach shown in Figure 5.

[0124] The methods and user devices for performing the hearing training described above have been explained separately with reference to Figures 3 to 5. However, it will be understood that these techniques can form alternative training modes within a broader range of methods. For example, the user device may be configured to perform any of the methods described in reference to Figures 3, 4, or 5 in response to user input, input from an external system, and / or in response to the execution of a standardized hearing test. For example, a standardized hearing test may identify a particular hearing training method as particularly beneficial to the user, and the user device may then be configured to perform that hearing training method. In this way, the hearing training can be easily customized to suit the user.

[0125] Furthermore, the techniques described above in relation to Figures 3 to 5 each include, but are not required, smartphones equipped with touchscreen displays 31, 41, and 51 (i.e., user devices 30, 40, and 50). In further examples, alternative user devices including devices and systems for providing AR and VR may be used. Similarly, in some examples of the present invention, hearing training may not include the display of a training environment. Instead, the training method may include the use of a physical training environment, or it may include only audio signals without the use of a visual training environment.

[0126] All of the devices and system components described above may be connected by wired or wireless connections.

[0127] Figure 6 shows a flowchart illustrating how the difficulty level of the listening training can be adapted to reflect the user's listening ability, increasing as the user's listening skills improve. This process can be performed using the user device 10 shown in Figure 1 and may involve any of the tasks described with reference to Figures 3 through 5.

[0128] In step s201, the method begins. In step s202, a hearing training session or round of hearing training, including multiple iterations of the hearing training method, is completed for a user (e.g., a single user). The repeated hearing training method may be the method described above with reference to Figure 2. The training session may include at least 10 iterations of a process in which a target audio signal and a background audio signal are provided to the user, user input is received, and the user input is analyzed to determine whether the user evaluation shown by the user input corresponds to information defined by the target audio signal.

[0129] Next, in step s203, the method includes determining the percentage of user evaluations represented by user inputs that precisely correspond to the information defined by each target audio signal throughout the training session. Optionally, feedback on the user's performance during the hearing training session is provided to the user based on the results of this determination. For example, the user may be provided with a feedback score in the form of a raw percentage or rating (e.g., number of stars).

[0130] Based on the results of this assessment in step s203, the difficulty level of the training is adapted or adjusted in steps s204 to s210, as described above. After adjusting the difficulty level of the hearing training, a new hearing training session (i.e., a new round of hearing training) may be started in step s211, and the process may be repeated.

[0131] Difficulty adjustment begins in step s204, where it is determined whether the percentage (e.g., the percentage of user successful iterations) is greater than a first threshold. If the percentage is greater than this first threshold, the training is considered too easy, and in step s205, the difficulty of future training sessions is increased. This first threshold may be in the range of 85% to 100%, preferably 90%.

[0132] If the percentage is less than the first threshold, the method proceeds to step s206 to determine whether the percentage is within the range from the first threshold to the second threshold. The second threshold may be in the range of 50-85%, preferably 80%. If the percentage is within this range, the training is determined to be appropriately difficult, and in step s207, the difficulty level of future training sessions is maintained at its existing level. Otherwise, the method proceeds to step s208.

[0133] In step s208, it is determined whether the percentage is less than the second threshold. If so, the hearing training is determined to be too difficult, and in step s209, the difficulty level of future training sessions is reduced. Otherwise, in step s210, the difficulty level is maintained at its existing level.

[0134] It should be noted that the step of determining whether the percentage in step s208 is less than the second threshold is an inherent consequence of the decisions in steps s204 and s206 not being met, and is therefore optional and redundant. Nevertheless, actively performing this step provides redundancy and can avoid errors and problems in the calculation process.

[0135] Therefore, the method shown in Figure 6 above involves periodically changing the difficulty level of the hearing training depending on whether a user meets a predetermined threshold for success or failure of user input over multiple consecutive iterations of the method. If the percentage of correct user input is greater than a predetermined first threshold, the difficulty level of the hearing training decreases in subsequent iterations of the hearing training process; or if the percentage of correct user input is less than a predetermined second threshold, the difficulty level of the hearing training increases in one or more subsequent iterations of the hearing training process.

[0136] Optionally, the process described above may be used to generate a baseline difficulty level for subsequent training sessions, but during subsequent training sessions, the difficulty level may be modified from this baseline level based on the user's performance during iterations of the hearing training method within the training session. Thus, the difficulty level is adjusted based on both previous training sessions and iterations of the method in the ongoing or current training session.

[0137] The method achieves improved training results when the first and second thresholds are set at 90% and 80%, respectively, such that the user's success rate is consistently maintained within the range of 80% and 90% (i.e., approximately 85%). This ensures that the tasks are presented to the user in a way that is challenging enough to keep them engaged but not so challenging that they frustrate them. Consequently, high user engagement is achieved, and users are more likely to continue listening training and significantly improve their listening skills.

[0138] Therefore, it will be understood that the method shown in Figure 6 can gradually increase or decrease the difficulty of the listening method to reflect changes in the user's listening skills. Changes in difficulty may include changes in the target audio signal, background audio signal, training environment, and the time scale to which the user must respond, as described above in the summary. Therefore, it will be understood that a wide variety of changes in the progression of difficulty between training sessions can be developed and predefined, depending on the desired progress of difficulty. For example, difficulty may be gradually increased by gradually decreasing the volume or quality of the target audio signal relative to the background audio signal, or by gradually increasing the number of sounds in the background audio signal and / or by changing their apparent positions. Similarly, these changes may be applied in combination or alternately as needed.

[0139] While the present invention has been described in the context of a fully functional data processing system, it is important to note that those skilled in the art will understand that the processes of the present invention can be distributed in the form of computer-readable media of instructions and in various other forms, and that the present invention applies equally regardless of the specific type of signal carrier medium actually used to perform the distribution.

[0140] Generally, any of the functions described in this text or shown in the figures can be implemented using software, firmware (e.g., fixed logic circuits), programmable or non-programmable hardware, or a combination of these implementations. The terms “component” or “function” as used herein generally refer to software, firmware, hardware, or a combination thereof. For example, in the case of a software implementation, the terms “component” or “function” may refer to program code that performs a specified task when executed on a processing unit. The separation of components and functions into separate units shown in the figures herein may reflect the actual physical grouping and assignment of such software and / or hardware, or may correspond to the conceptual assignment of different tasks performed by a single software program and / or hardware unit. Therefore, the various processes described herein can be implemented on the same processor or on any combination of different processors. For example, the analysis of user input, the generation of feedback in response to user input, the modification of a target audio signal or background audio signal, and / or any of the other processes described above may be performed by a processor within the user device, or by a processor in an external device or system (e.g., a remote or cloud-based system communicating with the user device).

[0141] The methods described above and the processes described herein can be materialized as code (e.g., software code) and / or data. Such code and data can be stored in one or more computer-readable media, which may include any device or medium capable of storing code and / or data used by a computer system. When a computer system reads and executes code and / or data stored in a computer-readable media, the computer system executes the methods and processes materialized as data structures and code stored in the computer-readable storage medium. In certain embodiments, one or more steps of the methods and processes described herein may be executed by a processor (e.g., a processor in a computer system or data storage system). Those skilled in the art will understand that computer-readable media include removable and non-removable structures / devices that can be used to store information such as computer-readable instructions, data structures, program modules, and other data used by a computing system / environment. Computer-readable media include, but are not limited to, volatile memory such as random access memory (RAM, DRAM, SRAM), non-volatile memory such as flash memory, various read-only memories (ROM, PROM, EPROM, EEPROM), magnetic memory and ferromagnetic / ferroelectric memory (MRAM, FeRAM), magnetic and optical memory (hard drives, magnetic tapes, CDs, DVDs), network devices, or other media currently known or to be developed that can store computer-readable information / data. Computer-readable media should not be understood or interpreted as containing propagated signals.

[0142] While specific embodiments of the Disclosure have been described, various modifications, changes, alternative structures, and equivalents are also included within the scope of the Disclosure. The embodiments of the Disclosure are not limited to operation within a particular data processing environment, but can freely operate within multiple data processing environments. Furthermore, while the embodiments of the Disclosure are described using a specific set of transactions and steps, it will be apparent to those skilled in the art that the scope of the Disclosure is not limited to the described set of transactions and steps. The various features and aspects of the embodiments described above may be used individually or in combination.

[0143] Therefore, the specification and drawings should be considered illustrative, not restrictive. However, it is clear that additions, deductions, deletions, and other modifications and alterations can be made without departing from the broader spirit and scope set forth in the claims. Thus, while specific disclosed embodiments have been described, they are not intended to be limiting. Various modifications and equivalents are included in the following claims. Modifications and alterations may include any relevant combination of the disclosed features.

Claims

1. A method for hearing training, which is performed using a user device that is a computer equipped with a user interface and audio output, The method involves providing a background audio signal and a target audio signal using the aforementioned audio output, wherein the target audio signal at least partially overlaps with the background audio signal, and the target audio signal defines information to be determined by the user. One or both of the background audio signal and the target audio signal include binaural audio. The user interface receives user input corresponding to the user evaluation of the information defined by the target audio signal, Based on the user evaluation indicated by the user input, provide feedback to the user, and repeat this process over and over again. The information defined by the target audio signal is, The content of the target audio signal or the linguistic content of the speech within the target audio signal, and / or, Including similarity to and / or relationship to a second target audio signal, Based on the determination that the user evaluation indicated by the user input corresponds to the information defined by the target audio signal, the difficulty level of the listening training is increased for subsequent iterations, and / or Based on the determination that the user evaluation indicated by the user input does not correspond to the information defined by the target audio signal, the difficulty level of the hearing training is reduced for subsequent repetitions. Methods for listening training.

2. The information defined by the target audio signal includes the target location in the training environment. Providing the aforementioned target audio signal means The user interface receives one or more auxiliary user inputs corresponding to intermediate positions within the training environment, The method involves changing one or more attributes of the target audio signal based on the relative positions of the intermediate position and the target position within the training environment. A method for hearing training according to claim 1, including the following:

3. Changing one or more attributes of the target audio signal is: The volume of the target audio signal related to the background audio signal is changed, The content of the target audio signal is changed, To change the pitch, duration, reverb, and / or rhythm of the target audio signal, or, When the target audio signal is binaural, the apparent sound source position of the target audio signal related to the user is changed. A method for hearing training according to claim 2, comprising one or more of the following.

4. The method for hearing training according to claim 1, wherein the information defined by the target audio signal corresponds to a target visual component in a plurality of different visual components in a training environment.

5. To provide two or more target audio signals using the audio output, It has, The method for hearing training according to claim 1, wherein the user input indicates whether the user believes the two or more target audio signals are similar and / or related.

6. The method for hearing training according to any one of claims 1 to 5, further comprising analyzing the user input received in the user interface and determining whether the user evaluation indicated by the user input corresponds to the information defined by the target audio signal.

7. The method for hearing training according to claim 1, wherein the difficulty level of the hearing training changes periodically over a series of iterations of the method, depending on whether the user meets a predetermined threshold for successful or unsuccessful user input.

8. Determining the percentage of user ratings indicated by user inputs that correctly correspond to the information defined by each of the target audio signals over a series of iterations of the method, It further possesses, A method for hearing training according to claim 1 or 7, wherein if the proportion of correct user input is greater than a predetermined first value, the difficulty level of the hearing training is reduced for one or more subsequent iterations, or if the proportion of correct user input is less than a predetermined second value, the difficulty level of the hearing training is increased for one or more subsequent iterations.

9. This includes a preliminary step of conducting a standardized hearing test. Based on the results of the standardized hearing test, The difficulty level of the aforementioned hearing training has been changed. The content, and / or one or more attributes of the target audio signal, and / or the background audio signal is changed, and / or The mode of the listening training will be changed. A method for hearing training according to claim 1.

10. The user device includes a display, The system includes displaying the training environment to the user using the aforementioned display. A method for hearing training according to claim 1.

11. The method for hearing training according to claim 1, wherein the target audio signal includes one or more of the following: human speech, animal sounds, traffic noise, musical instruments, weather noise, underwater noise, natural sounds, synthesized sounds, ambient noise, white noise, or synthesized sound effects.

12. The method for hearing training according to claim 1, wherein the background audio signal includes one or more of the following: human speech, animal sounds, traffic noise, musical instruments, weather noise, underwater noise, natural sounds, synthesized sounds, ambient noise, white noise, or synthesized sound effects.

13. The aforementioned background audio signal includes multiple sounds that overlap in at least part, A method for hearing training according to claim 1 or 12.

14. The user device is a smartphone, tablet, laptop, personal computer, or AR and / or VR system. A method for hearing training according to claim 1.

15. A user device having a user interface and an audio output, The user device is configured to perform the method for hearing training described in claim 1.

16. A non-temporary computer-readable medium that stores instructions for a user device to perform the hearing training method described in claim 1, when read by a processor.

Citation Information

Patent Citations

  • Auditory function training and detecting system based on virtual reality

    CN112451831A

  • Valve seat for batterfly valve

    JP1984097369A

  • Auditory diagnosis and training system apparatus and method

    US20110313315A1

  • External device leveraged hearing assistance and noise suppression device, method and systems

    US20170230769A1

  • Device for improving bloodstream volume in hearing-area of brain, and virtual sound source used for the same

    WO2009102052A1