Apparatus and method for determining audio processing parameters

The apparatus and method dynamically adjust audio processing parameters using user-specific learning systems to enhance sound reproduction devices' adaptability and quality in diverse environments.

JP7721683B2Active Publication Date: 2025-08-12FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2023571527
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-17
Filing Date
2022-05-16
Publication Date
2025-08-12
Estimated Expiration
2042-05-16

AI Technical Summary

Technical Problem

Conventional sound reproduction devices fail to adaptively adjust to individual user preferences and environmental changes, requiring repeated manual adjustments and failing to provide optimal audio quality in diverse listening environments.

Method used

An apparatus and method that determines audio processing parameters in real-time using user-specific adjustments integrated into a learning system, employing neural networks and reinforcement learning to adapt coefficients based on user operations and environmental audio signals.

Benefits of technology

Provides personalized audio processing parameters that enhance user experience by adapting to individual preferences and environmental conditions, reducing the need for manual adjustments and improving audio quality in dynamic listening scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007721683000001
    Figure 0007721683000001
  • Figure 0007721683000002
    Figure 0007721683000002
  • Figure 0007721683000003
    Figure 0007721683000003
Patent Text Reader

Abstract

The present invention relates to an apparatus and method for determining audio processing parameters in response to at least one audio input signal.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] SUMMARY OF THE INVENTION Embodiments in accordance with the present invention relate to an apparatus and method for determining audio processing parameters in response to at least one audio input signal.

[0002] Embodiments according to the present invention relate to apparatus and methods with artificial intelligence that can analyze audio signals during user operation, for example in a sound reproduction device, and assign or combine user-specific settings.

[0003] Furthermore, some embodiments relate to a concept for determining audio processing parameters based on audio signals acquired during user operations. [Background technology]

[0004] The individual perception of sound and therefore the individual requirements for sound or euphony for the adjustment of sound reproduction devices differ according to the following criteria:

[0005] ·Individuality · Situational needs External conditions Sound perception varies from person to person. For example, a conversation in a room with many people is more difficult for some people than for others. In addition, the same adjustment of sound reproduction, if necessary, is perceived differently. Environmental parameters such as the hearing environment also significantly affect the control values for sound adjustment of sound reproduction devices.

[0006] Current sound reproduction devices provide certain sound adjustments that are not applied in an automated manner. Sound reproduction devices, such as headphones, headsets, or portable devices for hearing assistance, such as hearing aids, often only include volume control and an equalizer for sound adjustment. Sound adjustments, such as volume amplification or treble or bass adjustment, are performed once by the user. It has become apparent that these adjustments must be performed again for each further sound reproduction to obtain continuously good audio quality.

[0007] Not only does the conventional approach require repeated sound adjustment processes for different sound reproductions, but it has also been shown that sound reproduction devices do not adaptively adjust to changes in the auditory environment, such as environmental sounds. It has been shown that even relatively small changes in environmental noise can increase listening effort for speech understanding.

[0008] Furthermore, conventional thinking has revealed that sound adjustments can only be performed based on sound default settings predetermined by manufacturers. This has revealed that this does not necessarily correspond to the individual needs of users. Thus, while settings such as "music" exist, preferences in music and personal intentions when listening to music are not taken into account. For example, expectations regarding the sound experience for opera singing and techno music are very different. In the default settings for the listening program "music," manufacturers simply base the settings on general assumptions that likely do not meet the requirements for the sound experience of either opera singing or techno music, and therefore can only provide users with insufficient sound reproduction.

[0009] Current sound reproduction devices for hearing assistance, such as hearing aids, can cost thousands of euros depending on their features, and expectations for these devices are therefore high. Hearing aid fitting is typically performed under laboratory conditions, often with only two speakers, and with only a very limited number of sounds, such as sine waves, noise, and speech. Complex noise situations, such as at intersections, cannot be simulated in an audiology laboratory, leading to frustration for hearing aid users and making it difficult to use in everyday life.

[0010] In learning applications for sound reproduction, such as the Github application "liketohear-ai-pt," the situation-specific parameter changes of the hearing aid algorithm, recorded by the user in a file, and the recorded frequency spectrum analysis assigned to the situation are processed by a self-learning algorithm. The algorithm establishes the relevance of specific frequency spectrums related to the user's decisions and automatically selects the assigned parameters as the basis for a predictive model. In a second step, the predictive model is applied to the previously recorded frequency spectrum analysis. It has become clear that the complexity of the frequency spectrum cannot be mapped by this learning application for sound reproduction, so further user adjustments are continually required.

[0011] In view of the above, what is needed is a concept for determining audio processing parameters at runtime that provides an improved trade-off between user convenience, obtainable audio quality, and implementation effort. Summary of the Invention

[0012] This object is solved by the subject matter of the independent claims. A central concept of embodiments of the present invention is the discovery of taking sound adjustments that are intuitively performed by the user at runtime and integrating them into a learning system in real time.

[0013] An embodiment according to the present invention includes an apparatus for determining audio processing parameters, such as parameters for audio processing, in response to at least one audio input signal, e.g., resulting from an audio input, the apparatus being configured to determine at least one coefficient of a processing parameter determination rule in a user-specific manner based on an audio signal obtained during a user operation, the apparatus being configured to obtain the audio processing parameters by using the processing parameter determination rule based on the audio input signal. The coefficients of the processing parameter determination rule may, for example, be coefficients of a neural network that receives the audio input signal or an input signal parameter extracted therefrom as an input quantity and provides the audio processing parameter as an output quantity. In other words, the coefficients of the processing parameter determination rule may be determined in a user-specific manner based on an input audio signal obtained during a user operation, e.g., during a user operation. Furthermore, the apparatus may be configured to obtain the audio processing parameters by using a processing parameter determination rule defined by at least one coefficient based on, for example, the audio input signal.

[0014] This embodiment is based on the central idea that per-user adjustment of one or more coefficients of a processing parameter determination rule based on an audio signal acquired during user operation allows the processing parameter determination rule to be adapted to the user's individual habits and needs. By using an audio signal acquired during user operation for per-user adjustment of the coefficients of a processing parameter determination rule, the coefficients can be better adapted to the (specific) hearing situation in which the user usually finds himself or herself. Thus, it is no longer necessary to pre-classify the acoustic environment (e.g., into a general category "music" and a general category "speech"), but the coefficients can also be adapted to the actual listening environment in which the user listens, for example, to music or speech, and to the user's individual needs. For example, by appropriate selection of the coefficients of the processing parameter determination rule, immediate per-user determination of audio processing parameters is possible; for example, the processing parameter determination rule adjusted by the coefficients needs to immediately determine the audio processing parameters without classifying the acoustic environment into one or several statically predetermined categories. Rather, the coefficients of the processing parameter decision rules can be adjusted based on the audio signal obtained during user operation, so that listening environments associated with users for which the user desires different audio processing parameters can be distinguished in a "hard" or "soft" manner (e.g., with smooth transitions).

[0015] In this way, by taking into account the audio signal acquired during user operation (and by respective adjustment of the coefficients of the processing parameter decision rules), the inventive concept makes it possible to provide very different audio processing parameters when, for example, speech is present in different acoustic environments in which the user is located (e.g., a busy open-plan office, a single office, an intersection with many trucks, an intersection with a tram, etc.), with the provided parameters typically being adapted to the settings desired by the user in each situation.

[0016] In this way, the inventive concept provides, with reasonable effort, audio processing parameters that suit the everyday realities and particular preferences of individual users.

[0017] According to a further embodiment, the device is configured to determine a database in response to user parameters adjusted by the user, such that entries in the database represent user parameters adjusted by the user. For example, the database can be established in real time during user operation, and a predictive model can be determined. Furthermore, the database, in that it includes information about the user parameters, can be used to determine coefficients of processing parameter determination rules. For example, the database can further include person-related control settings that can be linked to the user parameters. The user parameters adjusted by the user can replace audio processing parameters, for example, as output quantities, or can modify audio processing parameters, such that entries in the database represent user parameters adjusted by the user. For example, the database can be at least partially integrated into reinforcement learning accordingly, for example, using the user parameters adjusted by the user.

[0018] By establishing a database whose entries represent user parameters adjusted by a user, the coefficients of the processing parameter determination rules can be continuously improved or optimized, for example. The user parameters adjusted by a user (typically in different acoustic environments), which can be formed into a database and stored, for example, in a data bank or another memory structure, can represent audio processing parameter settings. For example, if there is an assignment of user parameters to audio signals (or audio signal characteristics) for each acoustic environment in which the user selected user parameters, this database can be used to determine the coefficients of the processing parameter determination rules. For example, by determining an increasingly larger database over time, for example, with a longer period of use by the user, a database that grows over time can exist for the (automatic) determination (or improvement) of the coefficients of the processing parameter determination rules, thereby enabling the increasing refinement or improvement of the coefficients (e.g., based on an increasing basis of different listening environments in which the user was located). In this way, the establishment and continuous expansion of the database can continuously improve the user experience.

[0019] According to a further embodiment, the device is configured to determine a database in response to at least one audio input signal, such that entries in the database represent the audio input signal. For example, the database can be used to determine coefficients of processing parameter determination rules. That is, for example, first, human-related control adjustments, for example, user parameters adjusted by a user, are stored and then augmented with sound information of the auditory environment as external conditions. This can generate data basis for providing coefficients of the processing parameter determination rules, for example, by using reinforcement learning.

[0020] According to a further embodiment, the device is configured to determine a database representing an assignment between different audio input signals and respective user parameters adjusted by the user. In other words, the device can assign external conditions to each other, for example, based on the audio input signal and human-related control settings, e.g., user parameters adjusted by the user. This means that the assignment can serve as the basis for, for example, a predictive model, which can be modified ad hoc based on the user's further sound adjustments, for example, by integrating the respective user parameters adjusted by the user with the database (which can then, for example, redetermine or improve the coefficients of the processing parameter determination rules). For example, in the background, an auditory scene can be continuously recorded and / or analyzed and / or evaluated by a microphone via audio input, and an analysis of the auditory scene can be generated, for example, via dynamics and / or frequency and / or spectral characteristics. The results of the auditory scene analysis can be integrated into the database, for example, as environmental parameters, and can be assigned to the user parameters to obtain a link between the user parameters and the audio input signal and the auditory environment at each time.

[0021] According to a further embodiment, the device is configured to determine a database, for example for determining coefficients of processing parameter decision rules, in response to an audio input signal, such that entries of the database describe or represent an audio output signal. By determining the database in response to at least one audio input signal and at least one audio output signal, for example a processing parameter decision rule of reinforcement learning can use the database to determine coefficients of processing parameter decision rules, for example for a neural network. The coefficients of the processing parameter decision rules can be obtained, for example, by common processing of the audio input signal and the assigned output signal, or by comparing the audio output signal with the audio input signal.

[0022] According to a further embodiment, the device is configured to determine a database, the database representing an assignment between different audio output signals and respective user parameters adjusted by a user. In other words, the database describes the assignment between different audio input signals, different audio output signals, and respective user parameters adjusted by a user, so that coefficients of processing parameter determination rules can be determined. The established database, for example, by analyzing input and output audio signals, can integrate sound processing into the training of a self-reinforcement learning algorithm. For example, the input audio signal or the audio input signal can include a sound environment, for example, an auditory environment. In other words, the established database, for example, by analyzing input and output audio signals, can select coefficients of processing parameter determination rules so that the desired connection between the audio input signal and the audio output signal is at least approximately achieved by the processing parameter determination rules.

[0023] According to a further embodiment, the device is configured to adjust at least one coefficient of the processing parameter determination rule based on the database obtained by the device to adjust the processing parameter determination rule in a user-specific manner to obtain audio processing parameters adjusted in a user-specific manner. In other words, for example, a reinforcement learning user model is adjusted based on artificial intelligence to obtain audio processing parameters adjusted in a user-specific manner or an audio signal adjusted in a user-specific manner. For example, it is possible to essentially learn and adapt to changes in the sound environment, e.g., the auditory environment, and user adjustments in a runtime manner. For example, the audio processing parameters adjusted in a user-specific manner can enable an audio signal added in a user-specific manner to be obtained during user operation when processing an audio input signal by using the audio processing parameters. In other words, a user-specific parameter set for sound processing can be obtained from a database or developed, which, on the one hand, applies the same control parameters in an automated manner under the same external conditions, but also allows further user adjustments in the situation itself, and these are integrated into the device as a learning system. For example, the learning system and application can adapt themselves to the user's sound preferences in a continuous learning process.

[0024] According to a further embodiment, the device is configured to provide and / or adjust the processing parameter decision rules based on a database, for example, the device can use the database to provide the processing parameter decision rules, e.g., by using reinforcement learning, in order to obtain an audio signal that is adjusted in a user-specific manner, e.g., by using audio processing parameters during user operation.

[0025] According to a further embodiment, the device is configured to determine and / or adjust at least one coefficient of the processing parameter determination rule based on at least one audio processing parameter corrected and / or compensated by a user. As already mentioned, the device can be configured to take into account or adjust user adjustments of user parameters during user operation and to allow further user adjustments of user parameters, for example, at a later time and therefore at the same location or therefore in the same sound environment, so that the previous user parameters are adjusted and / or overwritten with the newly adjusted user parameters. In other words, the coefficients of the processing parameter determination rule can be corrected by the user and / or the compensated audio processing parameters can be determined depending on, for example, the sound environment in which the user is located at each time.

[0026] According to a further embodiment, the device is configured to perform audio processing, e.g., parameterized audio processing rules, based on the audio input signal and on the audio processing parameters to obtain an audio signal that is adjusted in a user-specific manner, e.g., by taking into account user modifications of the audio processing parameters. In other words, the device can provide an audio signal for an audio output that is adjusted in a user-specific manner by optional audio processing of the audio input signal and the audio processing parameters. For example, the audio processing can be integrated into the device, resulting in an efficient system. Optionally, the audio processing can also be incorporated into the determination of the audio processing parameters.

[0027] According to a further embodiment, the device is configured to determine the coefficients of the processing parameter decision rules by using a comparison between the audio input signal and the audio input signal provided by using the audio processing parameters, for example by taking into account user modifications of the audio processing parameters. In other words, the determination of the coefficients of the processing parameter decision rules can be based on a comparison between the audio input signal and the direct audio output signal or the audio output signal provided by audio processing. For example, optionally, an audio analysis of the audio input signal or an audio analysis of the audio output signal can be performed before or after using the comparison, and the coefficients of the comparison parameter decision rules can be determined based on the audio analysis results of the audio signal. Determining the coefficients of the parameter decision rules by using such a comparison provides particularly reliable or robust results, since the audio signal actually output to the user can be used as the basis for determining the coefficients of the parameter decision rules. The criterion that the audio output signal should correspond to the audio output signal desired by the user is more significant and robust than pure optimization of the audio processing parameters themselves.

[0028] According to a further embodiment, the device is configured to provide user parameters adjusted by the user as output quantities instead of audio processing parameters, where the user parameters adjusted by the user include volume parameters and / or sound parameters and / or equalizer parameters. In other words, the user parameters may include, for example, sound design and / or filter parameters for equalizing sound frequencies. Providing the user parameters adjusted by the user as output quantities, for example, allows for immediate user intervention, resulting in a significantly better user experience. The user intervention can further be used to improve the coefficients, possibly preventing future user intervention (and instead automatically obtaining adjustments that are adapted to the user's requirements).

[0029] According to a further embodiment, the device is configured to, for example additionally, combine user parameters with the audio processing parameters to obtain a composite parameter of the audio processing, which is provided as an output quantity. The composite parameter may, for example, comprise user parameters and audio processing parameters that are provided to the audio processing in a combined manner or that are combined by using the audio processing and provided as an output quantity, for example, to reinforcement learning. Thus, a fast user intervention is possible and the audio processing can be adapted to the user requirements.

[0030] According to a further embodiment, the device is configured to perform an audio analysis of the audio input signal to provide an audio input signal analysis result for determining at least one coefficient of the processing parameter determination rule, for example, by using a processing parameter determination rule. For example, the processing parameter determination rule may define a derivation rule for deriving an audio processing parameter from the audio input signal analysis result. The audio analysis of the audio input signal may provide the audio input signal analysis result, for example, in the form of information about the spectral characteristics and / or dynamics and / or frequency of the audio input signal, or as information about intensity values per band. The audio input signal analysis result may be provided as an input quantity for determining one or more coefficients of the processing parameter determination rule, for example, by using reinforcement learning. Note that in some embodiments, the audio analysis may analyze and evaluate an audio input signal resulting from the audio input in advance to provide the processing parameter determination rule, but this is not required. For example, additional information about the spectral characteristics of the audio input signal may be obtained as the audio input signal analysis result. Furthermore, by using the audio input signal analysis result, the processing parameter determination rule may be configured in a simpler manner, for example, compared to using the complete audio input signal to determine the audio processing parameters. In this way, the parameters or values of the audio input signal analysis result can describe the essential characteristics of the audio input signal in an efficient way, such that the processing parameter determination rule involves a relatively small number of input variable (i.e., for example) parameters or values of the audio input signal analysis result and can therefore be implemented in a relatively simple manner. In this way, good results can be obtained with little effort.

[0031] According to a further embodiment, the device is configured to perform an audio analysis of the audio input signal, for example by using a processing parameter determination rule, to provide an audio output signal analysis result in the form of information about spectral characteristics of the audio input signal, for example for determining at least one coefficient of the processing parameter determination rule. In other words, the device is configured to perform an audio analysis before the processing parameter determination rule or after the processing parameter determination rule to provide either an audio input analysis signal result or an audio output signal analysis result, or both, for determining the coefficients of the processing parameter determination rule. For example, by determining the audio output signal analysis result, it is very easy to compare the audio input signal with the audio output signal, and for example, the values or parameters of the audio output signal analysis result can describe the characteristics of the audio output signal very efficiently (or in a very compact form). Therefore, the coefficients of the processing parameter determination rule can be determined or optimized in a very efficient manner, and the user's desired processing can be achieved, for example, by evaluating the audio output signal analysis result in an efficient manner, or a comparison between the audio input signal analysis result and the audio output signal analysis result can enable a conclusion regarding the coefficients of the processing parameter determination rule.

[0032] According to further embodiments, the audio processing parameters include at least one multiband compression parameter R and / or at least one hearing threshold adjustment parameter T and / or at least one band-dependent amplification parameter G and / or at least one interference noise reduction parameter and / or at least one blind source separation parameter. Furthermore, the audio processing parameters may include at least one sound direction parameter and / or binaural parameters and / or parameters related to the number of different speakers and / or parameters for adaptive filters in general, such as hall suppression, feedback, echo cancellation, and active noise cancellation (ANC). For example, the sound direction parameter may select or adjust the directionality of a sound source so that, for a given combination of audio processing parameters, only sounds from a desired direction, e.g., a conversation partner, are processed. It has been shown that such audio processing parameters can affect audio signal processing in an efficient manner, and that it is already possible to affect audio signal processing over a wide adjustment range with a small number of parameters that can be easily determined by processing parameter determination rules.

[0033] According to a further embodiment, the device may include a neural network that executes a processing parameter determination rule, for example, so that at least one coefficient is defined, or preferably multiple coefficients are defined, and is configured to obtain audio processing parameters by using the processing parameter determination rule. Furthermore, the neural network may be configured to obtain audio processing parameters based on an audio input signal, either directly from the audio input or as an audio input signal analyzed by interconnected audio analysis. It has been shown that neural networks are well suited to determining audio processing parameters, and that the coefficients can be well adapted to the personal perception of individual users. For example, a neural network, in which edge weights can be defined by coefficients of the processing parameter determination rule, can be adapted to the user's needs by coefficient selection (e.g., by a training rule). The coefficients can be continuously improved, for example, in the presence of further user adjustments. This can result in a significantly better user experience.

[0034] According to further embodiments, the device is configured to provide and / or adjust processing parameter determination rules based on reinforcement learning methods, unsupervised learning methods, multivariate prediction methods, and / or a multidimensional parameter space determined by multivariate regression to determine the audio processing parameters. The processing parameter determination rules can, for example, provide coefficients of a neural network based on reinforcement learning methods. The multivariate prediction methods can, for example, include prediction of frequency bands and / or prediction of input / output characteristics depending on user parameters. Furthermore, a multivariate regression method can, for example, analyze all existing frequency bands to determine the multidimensional parameter space. The multidimensional parameter space can, for example, be a two-dimensional parameter setting with a graphical surface whose axes include volume and sound adjustments or on which the user can adjust or continuously adjust the user parameters by means of sliders or points on a coordinate system assigned to the volume and sound adjustments. The above-described method enables the device to determine audio processing parameters such that, for example, a learning algorithm adjusts audio processing parameters for each user, e.g., such that the audio processing parameters provided by applying the processing parameter determination rules approach the audio processing parameters corrected by the user as the learning progresses, e.g., such that the processing parameter determination rules adjust themselves in a continuous learning process, e.g., in response to user adjustments of the audio processing parameters. As expected, for example, the method's access to a database or data memory is unlimited (e.g., so that as the size of the database increases, even better coefficients can be determined by using the above-described learning method).

[0035] According to a further embodiment, the device is configured to obtain user parameters adjusted by the user via or by an intuitive and / or ergonomic user control, such as an interface, for example, a user interface, a 2D space on a smartphone display, etc. In other words, the device can include an interface (for example, an electrical interface or a man-machine interface) for adjusting the user parameters. Preferably, the visual user control can include, for example, a volume adjustment with a slider for adjusting louder and softer volume and / or height and depth. In this way, parameter adjustment can be made extremely easy for the user, and this simple sound adjustment has been shown to often already result in a good auditory impression.

[0036] According to further embodiments, the audio input signal includes a multi-channel audio signal, e.g., having at least four channels or at least two audio channels. For example, the audio input signal can be provided by an audio input, e.g., from, through, or by a microphone. Furthermore, the audio input signal can include information such as the number of channels and / or the number of frequency bands. The use of a multi-channel signal allows, for example, localization of a desired sound source and / or a disturbing sound source, as well as consideration of the direction of the desired sound source or the disturbing sound source, when determining the coefficients of an audio processing parameter or a processing parameter decision rule.

[0037] According to a further embodiment, the device is configured to perform audio processing separately for at least four frequency bands of the audio input signal. In this way, it can be ensured that frequency selectivity is provided so that each individual frequency can be analyzed, for example when the audio input signal comprises a multi-channel audio signal. By taking into account different intensities in different frequency bands, different acoustic environments can be taken into account, and also specific user requirements regarding frequency response can be efficiently taken into account.

[0038] According to a further embodiment, the device is configured to acquire audio processing parameters in real time, e.g., at run time during user operation, and / or determine at least one coefficient of a processing parameter determination rule in a user-specific manner, e.g., continuously, sequentially, during user operation, e.g., in real time, to determine and / or adjust corrected audio processing parameters in real time. In other words, the device is configured to determine and / or adjust audio processing parameters in real time, e.g., such that the device, as a learning system, performs this learning process in real time, e.g., during user operation. In other words, in the present invention, sound processing is controlled based on external conditions, e.g., measured in real time. Therefore, the analysis of all existing frequency bands is also performed in real time so that a predictive model can be provided based on real-time multidimensional optimization, which means, for example, optimization in which audio processing parameters are determined based on the analyzed frequency bands and user parameters stored in a data memory.

[0039] According to a further embodiment, the present invention comprises a hearing aid, the hearing aid comprising audio processing, the hearing aid comprising a device for determining audio processing parameters, the audio processing being configured to process an audio input signal in response to the audio processing parameters. For example, the hearing aid may implement or integrate a device for improving the individual perception of sounds or tones in the form of an audio signal for a user. It has been shown that the devices described herein are particularly well suited for use in hearing aids, and that the hearing impression can be significantly improved by using the concepts of the present invention.

[0040] An embodiment according to the invention comprises a method for determining audio processing parameters in response to at least one audio input signal, the method comprising: determining, in a user-specific manner, at least one coefficient of a processing parameter determination rule based on an audio signal acquired during a user operation; and obtaining audio processing parameters by using the processing parameter determination rule based on the audio input signal. The method is based on the same considerations as the above-described apparatus and can optionally be supplemented by all the features, functions and details described herein for the inventive apparatus. The method can be supplemented by the above-described features, functions and details, both individually and in combination.

[0041] A further embodiment according to the invention comprises a computer program having a program code for performing the method when the computer program runs on a computer.

[0042] In the following, embodiments will be described with reference to the accompanying drawings. [Brief explanation of the drawings]

[0043] [Figure 1] 1 shows a schematic block diagram of an apparatus according to an embodiment for determining audio processing parameters in response to at least one audio input signal; [Figure 2]1 shows a schematic block diagram of an apparatus according to one embodiment for determining audio processing parameters in response to at least one audio input signal by reinforcement learning based on an audio input signal and an audio output signal. [Figure 3] 1 shows a schematic block diagram of an apparatus according to one embodiment for determining audio processing parameters in response to at least one audio input signal by reinforcement learning based on an audio analysis of the audio input signal and an audio analysis of the audio output signal. [Figure 4] 1 shows a schematic block diagram of an apparatus according to one embodiment for determining audio processing parameters in response to at least one audio input signal by audio analysis of the audio input signal and reinforcement learning based on user parameters adjusted by a user; [Figure 5] 1 shows a schematic block diagram of an apparatus according to one embodiment for determining audio processing parameters in response to at least one audio input signal by reinforcement learning based on the audio input signal and user parameters adjusted by a user. [Figure 6] 1 shows a schematic flow diagram of a method according to one embodiment for determining audio processing parameters; DETAILED DESCRIPTION OF THE INVENTION

[0044] Before describing the embodiments of the present invention in more detail based on the drawings, it should be noted that identical, functionally equivalent or comparable elements, objects and / or structures are given the same reference numerals in different figures, and therefore the descriptions of these elements shown in different embodiments are interchangeable or mutually applicable.

[0045] The embodiments described below are described in the context of multiple details. However, the embodiments may be practiced without these detailed features. Furthermore, the embodiments are described by using block diagrams instead of detailed representations to facilitate understanding. Furthermore, details and / or features of individual embodiments may be readily combined with one another unless explicitly stated to the contrary.

[0046] 1 shows a schematic block diagram of an apparatus 100 for determining audio processing parameters 120, where the audio processing parameters 120, shown at an output side of the apparatus 100, are determined in response to at least one audio input signal 110, shown at an input side of the apparatus 100. The exemplary schematic diagram of the apparatus 100 includes, for example, coefficient determination, indicated by a coefficient determination 130 block, where coefficients 132 of the coefficient determination 130 can be provided to processing parameter determination rules 140. The audio input signal 110 can, for example, be used directly by the processing parameter determination rules 140 to obtain the coefficients 142 of the processing parameter determination rules 140, and / or can be used as an audio signal 112 obtained during user operation by the coefficient determination 130 to provide the coefficients 132 to the coefficient determination 130. For example, the coefficient determination 130 can be performed in a user-specific manner during user operation, such that the coefficients 132 are provided to the coefficient determination 130 of the processing parameter determination rules 140, and the audio processing parameters 120 are obtained by using the processing determination rules 140 based on the audio input signal 110.

[0047] Thus, the coefficients of the processing parameter determination rules can be adjusted, for example, so that the processing parameter determination rules are based on an audio input signal and use the coefficients to provide as output audio processing parameters that, when used in audio processing, result in an audio output signal that meets user expectations.

[0048] 2 shows a schematic block diagram of an apparatus 200 according to one embodiment. The illustrated apparatus 200 for determining audio processing parameters includes, for example, an audio input unit 210, an audio processing unit 220, user controls 230, an audio output unit 240, and processing decision rules (or processing parameter determiners) in the form of a reinforcement learning unit 250 and a neural network 260.

[0049] The audio input unit 210 may include, for example, a microphone or other audio detection device and may include, for example, information regarding the number of channels, such as "C," and / or the number of frequency bands, such as "B." For example, tones, sounds, or sound waves, or generally audio signals, may be received via the audio input unit 210 and provided, for example, as audio input signals 212, 214, and 216 for the audio processing unit 220 and / or the reinforcement learning unit 250 and / or the neural network 260. For example, the audio signal 212 for the neural network 260, the audio signal 214 for the reinforcement learning unit 250, and the audio signal 216 for the audio processing unit 220 may be provided, where the audio signals 212, 214, and 216 may be the same or different in details (e.g., with respect to sampling rate, frequency resolution, bandwidth, etc.). Thus, here, audio signal 212 may be equal to audio signal 214 and / or audio signal 216 (or may at least describe the same audio content) and may have the same corresponding information regarding the number of frequency channels and frequency bands, and therefore the audio signal may be split, for example, directly by audio input unit 210 without the need for further audio analysis and provided, for example, via several output units or data paths of audio input unit 210.

[0050] The audio processing unit 220 may, for example, include one and / or several parameterized audio processing rules that process one or several audio signals 216, and provide, for example, an audio signal 217 adjusted in a per-user manner (or provide several audio signals adjusted in a per-user manner) based on the input audio signal 216 (or multiple input audio signals), for example by using parameterized audio processing rules parameterized by composite parameters 272. The audio processing unit 220 uses the composite parameters 272, for example by using parameterized audio processing rules, to process the audio input signal 216 based on the audio input 210, and obtain the audio signal 210 adjusted in a per-user manner. Optional details and embodiments regarding the composite parameters 272 are described in more detail below in the present patent application. Before that, further details and embodiments regarding the components of the device 200 will be described.

[0051] The audio output unit 240 can, for example, receive the corrected and newly assigned audio signal 217 adjusted in a per-user manner and provide it to a coefficient determination unit 250, for example realized by using reinforcement learning, as a corrected or processed audio signal 218 for determining parameters or coefficients of a processing parameter determination rule (e.g., a neural network 260). Alternatively or additionally, the audio output unit can, for example, provide the corrected, newly assigned and adjusted audio signal 217 adjusted in a per-user manner by the audio processing unit 220 as a corrected or processed audio signal 219 for, for example, an interface for headphones or speakers, although this is not required.

[0052] Additionally, some embodiments allow for providing additional information about the audio signal 218 to the reinforcement learning unit 250 (or another means for determining coefficients or parameters of a processing parameter determination rule) via the audio output unit 240, for example to provide information about the audio signal to a data store 252 (the contents of which may be part of a database).

[0053] Similar to the audio input signal 214, the audio output signal 218 can be provided to a reinforcement learning unit 250, for example, to determine coefficients or parameters of a processing parameter determination rule 260, and information of the audio input signal 214 and the audio output signal 218 is stored in a data memory 252 as respective databases of the device 200, for example.

[0054] In other words, for example, the reinforcement learning unit 250 can determine coefficients or parameters of the processing parameter determination rules 260 according to the audio signals 218 and 214. Furthermore, the reinforcement learning unit 250 can, for example, augment a database based on the audio signals 214, 218 and / or incorporate the audio signals 214, 218 into the data store 252. Alternatively or additionally, the reinforcement learning unit can determine or store at least one user adjustment coefficient 254 in the database.

[0055] However, it should be noted that use of the output audio signal 218 by the reinforcement learning unit 250 (or another device for determining coefficients of processing parameter decision rules that can replace the reinforcement learning unit 250) is considered optional.

[0056] The database or data store 252 may contain a plurality of pieces of information, for example, information about the audio input unit 210 (or audio input signal) and / or one or more of the audio signals 212 and 214 resulting from the audio input unit 210, and / or information about the audio output unit 240 and / or the audio signal 218 resulting from the audio output unit 240, and / or information about the audio processing unit 220, and may also include, for example, at least one user adjustment coefficient 254. The user adjustment coefficient 254 may be a coefficient determined for use by the processing parameter determination rules 250, for example, based on the database 252 and / or based on the adjusted user parameters 232. The user adjustment coefficient may be an audio processing parameter that has been adjusted by a user.

[0057] The coefficients of the processing parameter decision rules, i.e., for example, the edge weights of a neural network, may be based, among other things, on a method of reinforcement learning, shown in FIG. 2 as "reinforcement learning unit" at reference number 252.

[0058] For example, the reinforcement learning unit 250 (e.g., as a subfunction) can determine the database or contents of the data store 252 such that the data store 252 describes the assignments between different audio input signals 212, 214 and respective user parameters 232 adjusted by the user, e.g., user adjustment coefficients 254.

[0059] The neural network coefficients 256 may be provided by the reinforcement learning unit 250 in an advantageous manner in that the reinforcement learning unit 250 determines the database or contents of the data store 252 such that the data store 252 (e.g., additionally) describes the assignment between the audio output signal 218 and each user parameter adjusted by the user, e.g., user adjustment coefficients 254.

[0060] Then, the processing parameter determination rules can be configured as or integrated into a neural network 260, for example, to obtain audio processing parameters 262 by using the coefficients 256 determined by the reinforcement learning unit 250. In other words, for example, the neural network 260 can determine the audio processing parameters 262 based on the audio signal 212 and the coefficients 256 obtained by the reinforcement learning unit 250, and as a result, for example, a learning algorithm adjusts the audio processing parameters 262 for each user.

[0061] The at least one audio processing parameter 262 provided by the neural network 260 may be a single parameter or may include several parameters. The neural network 260 may provide, for example, one or several of the following parameters as the audio processing parameter 262: a user profile parameter N and / or a multiband compression parameter R and / or an auditory threshold adjustment parameter T and / or smoothing (e.g., one or several smoothing parameters) and / or compression adjustment (or one or several compression parameters). Furthermore, for sound adjustment, one or more parameters such as (alternatively or additionally) band-dependent amplification G, interference noise reduction (or one or several interference noise reduction parameters), and / or blind source separation (or one or several parameters of blind source separation) may be used (or provided as audio processing parameter 262 by the neural network).

[0062] For example, the number of input parameters (e.g., of the reinforcement learning unit 250 and / or the neural network 260) may result in a dependency on the number C of channels of a multi-channel audio signal, and also on the number B of processing bands or the number P of user parameters. For example, the number P of user parameters may result as the product of the number B of frequency bands and the number C of audio signals or audio channels.

[0063] Alternatively or additionally, the input parameters (e.g., of a reinforcement learning unit or neural network) may include audio features N, e.g., every 10 ms, e.g., F=2048 Fourier coefficients per channel for the input (e.g., audio input signal) and output (e.g., audio output signal).

[0064] For example, the number of output parameters (e.g., output parameters of neural network 260 or input parameters of audio processing) in learned user profile M may be composed of the number of audio channels (e.g., C), hearing threshold adjustment T, multiband compression with rate R, band-dependent amplification G, and two additional time constants, where the number of values of G, R, and T corresponds, for example, to the number of bands B. Furthermore, the values of learned user profile M (or multiple values of learned user profile M) can form user adjustment coefficients (or parameters) 254 (or a set of user adjustment coefficients or parameters).

[0065] The user controls 230 provide at least one user parameter 232, which may include, for example, a volume parameter and / or a sound adjustment parameter. The user controls may include, for example, an interface for visualizing one or several user parameters.

[0066] Volume control or volume adjustment that can be performed by user controls 230 can provide parameters that, for example, result in amplification or attenuation of the audio signal. Depth adjustment, height adjustment, and / or equalizer allow a user to adjust sound adjustment parameters, for example, via user controls 230, which can be combined with audio processing parameters 262 (provided by neural network 260), for example, by using combiner 270 as part of user parameters 232.

[0067] In other words, the user parameters 232 provided by the user controls 230 may be combined with the audio processing parameters 262, for example, by addition, multiplication, division, or subtraction. A combination 270 of the user parameters 232 and the audio processing parameters 262 may provide, for example, a composite parameter 272 to the audio processing unit 220. Alternatively, the user parameters 232 may replace the parameters 262, for example, if the user desires a significantly different adjustment from the adjustment predetermined by the parameters 262.

[0068] In summary, the device 200 can be described as processing an audio input signal obtained via the audio input unit 210 in the audio processing unit 220 to adapt the sound characteristics to the user's desires or needs. The processing characteristics of the audio processing unit 220 are adjusted by parameters 272, which are governed on the one hand by the neural network 260 and on the other hand can be modified by the user via the user controls 230. In general, the reinforcement learning unit 250 serves to adjust one or several coefficients (e.g., edge weights) of the neural network so that the parameters provided by the neural network essentially correspond to the user's expectations, i.e., include the parameter values adjusted by the user via the user controls 230 in each different acoustic environment within an acceptable tolerance.

[0069] It is therefore possible that after sufficient training in a number of different acoustic environments, the device will arrive at automatic settings for audio processing that are agreeable to the user.

[0070] FIG. 3 shows a schematic or block diagram of an apparatus 300 for determining audio processing parameters in response to at least one audio input signal based on the apparatus 200 of FIG.

[0071] It should be noted that in the device 300 according to Fig. 3, the functional blocks also shown in Fig. 2 may have similar or equivalent functions to (but do not necessarily include) the respective functional blocks in the device 200, for example. It should further be noted that the device 300 may optionally be supplemented by all the features, functions and details described herein, either alone or in any combination.

[0072] Similar to device 200, device 300 has an audio input unit 310 (which may correspond to audio input unit 200), an audio processing unit 320 (which may correspond to audio processing unit 220), a user control unit 330 (which may correspond to user control unit 230), an audio output unit 340 (which may correspond to audio output unit 240), a reinforcement learning unit 350 (which may, for example, correspond to reinforcement learning unit 250 in terms of its basic functionality), a neural network 360 (which may, for example, correspond to neural network 260 in terms of its basic functionality), and a combination unit 370 (which may, for example, correspond to combination unit 270) of user parameters 332 adjusted in a user-specific manner with audio processing parameters 362.

[0073] Starting from the device 200 of FIG. 2, the device 300 of FIG. 3 further includes or comprises an audio analysis unit 380-1 between the audio input unit 310 and the neural network 360, and an audio analysis unit 380-2 between the audio output unit 340 and the reinforcement learning unit 350.

[0074] In particular, this configuration enables the audio analyzer 380-1 to receive and analyze the audio input signal 311, e.g., provided by the audio input unit 310, and provide an audio input signal analysis result, e.g., information about the spectral characteristics and / or dynamics and / or frequency of the audio input signal 311, in the form of audio analysis signals 312 and / or 314. Information about the audio analysis result of the audio analyzer 380-1 can be provided to the neural network 360 and the reinforcement learning unit 350 (e.g., simultaneously) by means of the analyzed audio signals 312, 314, for example.

[0075] For example, the processing parameter determination rules may include part of the neural network 360 (or part of the reinforcement learning unit 350) or may define derivation rules for deriving the audio processing parameters 362 from the audio input analysis results. The audio analyzer 380-1 may obtain additional (or compact) information about spectral characteristics, such as intensity values per frequency band and channel, to provide, for example, frequency selectivity of the audio signal (e.g., a multi-channel audio signal). Frequency selectivity is necessary to enable analysis and representation of perceptible sound aspects of the signal. In general, the audio analyzer 380-1 may significantly reduce the amount of input data to the neural network, compared to, for example, the idea of inputting time-domain sample values to the neural network. For example, the complexity of the neural network 360 can be kept relatively low in that the analyzed audio signals 312, 314 include parameters that describe the characteristics of the audio input signal in a compact form (e.g., the number of parameters per time portion is 10 times less, or 20 times less, or 50 times less than the number of samples per time unit). Thus, the number of coefficients of the neural network can be kept relatively low, thus facilitating the learning process (e.g., by the reinforcement learning unit 350). This is even more true the more suitable the parameters of the analyzed audio signal are for distinguishing between different acoustic environments.

[0076] Further, optionally, an audio analysis 380-2 of the audio output signal 342 can be performed to provide an audio output signal analysis result for determining at least one coefficient of a processing parameter rule, e.g., at least one coefficient of the reinforcement learning unit 350.

[0077] A "common" audio analysis of the audio input signal 311 and the audio output signal 342 (e.g., audio analysis of both the audio input signal and the audio output signal) is also possible, and separate audio signal analysis results can be provided. In this context, separate means that the audio input signal analysis results can be provided to a separate component than, for example, the audio output signal analysis results. For example, the information in the audio analyses 380-1, 380-2 of the input or output signals can be different from each other, or can be similar or identical accordingly.

[0078] Here, in some embodiments, audio output unit 340 provides corrected or processed audio signal 319 for interfacing, for example, to headphones or speakers, although this is not required. Additionally, some embodiments allow audio analysis 280-2 to provide audio signal 313 for interfacing or further interfacing. This allows device 300 to provide audio signals 319 and 313 via at least one interface, for example, for an external component, although this is not required.

[0079] In summary, it can be stated that in the device 300, one or several respective audio analysis results, rather than the input or output audio signals themselves, are fed to the neural network 360 or the reinforcement learning 350. Thus, by appropriate pre-analysis of the input and / or output audio signals, the complexity of the neural network and thus of the reinforcement learning can be kept low, which significantly reduces the implementation effort.

[0080] FIG. 4 shows a schematic block diagram of an apparatus 400 for determining audio processing parameters in response to at least one input signal that is based in part on the apparatus 200 of FIG.

[0081] It should be noted that in the device 400 according to Fig. 4, the functional blocks also shown in Fig. 2 may have similar or identical functions to (but do not necessarily include) the respective functional blocks in the device 200, for example. It should further be noted that the device 400 may optionally be supplemented by all the features, functions and details described herein, either alone or in any combination.

[0082] The device 400 includes an audio input unit 410 (which may correspond, for example, to the audio input unit 210), an audio processing unit 420 (which may correspond, for example, to the audio processing unit 220), a user control unit 430 (which may correspond, for example, to the user control unit 230), an audio output unit 440 (which may correspond, for example, to the audio output unit 240), a reinforcement learning unit 450 (which may correspond, for example, to the reinforcement learning unit 250 in terms of its basic functionality), a neural network 460 (which may correspond, for example, to the neural network 260 in terms of its basic functionality), a combination unit 470 (which may correspond, for example, to the combination unit 270), and an audio analysis unit 480 (which may correspond, for example, to the audio analysis unit 380-1) between the audio input unit 410 and the neural network 460 and the reinforcement learning unit 450.

[0083] Compared to apparatus 300, apparatus 400 does not include an audio analysis of audio output 440, and compared to apparatus 200, no audio output signal is provided from audio output 440 to reinforcement learner 450. In other words, reinforcement learner 450 does not receive information about the audio output signal.

[0084] Instead, the reinforcement learning unit 450 is based on information 433 describing composite parameters 472, 473 or modifications or adjustments of audio processing parameters 462 provided by the neural network 460 by a user. Additionally, the reinforcement learning unit uses the audio input signal analysis results 414.

[0085] In other words, the reinforcement learning unit 450 can determine the database 452 in response to the user parameters or composite parameters 472, 473 adjusted by the user, such that entries in the database 452 represent the user parameters 472, 473 adjusted by the user. The database 452 can be provided to or used to determine the processing parameter determination rules or coefficients 456 of the neural network 460. This allows for the determination of a predictive model that is directly based on the user parameters (or the audio signal processing parameters 472 adjusted by the user) directly assigned to the reinforcement learning unit 450.

[0086] Optionally, one or several composite parameters 472, 473 or user parameters can be directly introduced into the neural network via composite parameters 474, 460, for example, so that compression settings and / or other parameters can be provided for the audio processing parameters 462 as output.

[0087] Alternatively or optionally, each user parameter 432 adjusted by the user can be provided directly to the reinforcement learning unit 450 (as indicated by reference numeral 433), although this is not required. For example, information about how the user modifies the parameters 462 provided by the neural network 460 can be used for reinforcement learning. If the user does not modify the parameters 462 provided by the neural network 460 at all, or modifies them only slightly, it can be assumed that the user is completely satisfied, or at least to a large extent, with the current functioning of the neural network, and therefore no, or only a small, correction of the neural network coefficients is required. However, if the user performs significant changes to the parameters 462, the reinforcement learning unit can assume that significant changes to the neural network coefficients are required so that the neural network provides parameters 462 that correspond to the user's expectations. In this way, for example, information 433 describing the user's intervention can be used by the reinforcement learning unit to trigger learning and / or determine the extent of change of the neural network coefficients.

[0088] Overall, the embodiment according to FIG. 4 allows for efficient learning and / or (eg, continuous) improvement of the coefficients 456 of the neural network 460.

[0089] Figure 5 shows an apparatus 500 having similar characteristics to the apparatuses 200, 300 and 400. It should be noted that in the apparatus 500 according to Figure 5, the functional blocks also shown in Figures 2, 3 and 4 may (but do not necessarily have to) have similar or identical functions to the respective functional blocks in the apparatuses 200, 300 and 400, for example. It should further be noted that the apparatus 500 may optionally be supplemented by all the features, functions and details described herein, either alone or in any combination.

[0090] The schematic block diagram of FIG. 5 shows an apparatus 500 including an audio input unit 510 (which may correspond, for example, to audio input unit 210), an audio processing unit 520 (which may correspond, for example, to audio processing unit 220), a user control unit 530 (which may correspond, for example, to user control unit 230), an audio output unit 540 (which may correspond, for example, to audio output unit 240), a reinforcement learning unit 550 (which may correspond, for example, to reinforcement learning unit 250 in terms of its basic functionality), a neural network 560 (which may correspond, for example, to neural network 260 in terms of its basic functionality), and a combination unit 570 (which may correspond, for example, to combination unit 270).

[0091] Apparatus 500, for example, does not include audio analysis of the audio input signal and does not include audio analysis of the audio output signal, so that audio signals 512 and 514 can be routed directly from audio input unit 510 to reinforcement learning unit 550 or neural network 560. Optionally, audio analysis of the audio input signal can also be performed in apparatus 500.

[0092] 2 of apparatus 200, audio input signal 512 may be provided for neural network 560, and audio input signal 514 may be provided for reinforcement learning unit 550. In contrast to apparatus 400, reinforcement learning unit 550 of apparatus 500 may be based on audio input signal 514 and one or several audio processing parameters 572 that are provided to (or actually used by) audio processing unit 520.

[0093] Optionally, user parameters or composite parameters 572 can be provided to neural network 560 such that user parameters 572 and coefficients provided by reinforcement learning unit 550 are taken as inputs to neural network 560 or provided to neural network 560.

[0094] The apparatus 500 allows for highly efficient adjustment of the neural network coefficients, since the reinforcement learning unit 550 takes into account the parameters actually used by the audio signal processing unit 520 and can therefore determine or optimize the neural network coefficients very accurately.

[0095] 6 shows a schematic flow diagram of a method 600 for operating an apparatus such as apparatus 100, 200, 300, 400, or 500, or generally for obtaining audio processing parameters. A first step 610 involves determining at least one coefficient of a processing parameter determination rule for each user based on an audio signal obtained during user operation. A second step 620 involves obtaining audio processing parameters by using the processing parameter determination rule based on an audio input signal.

[0096] Here, method 600 is performed, for example, such that audio processing parameters are determined in response to at least one audio input signal. Here, method 600 can be performed such that sound processing or audio processing based on immediately recorded environmental sounds (e.g., the audio input signal results in adjustment of the audio processing parameters) results in an improvement of an individual's perception of sound. For example, coefficients of processing parameter determination rules can be determined in a user-specific manner (e.g., in real time) based on audio input signals obtained during user operation, such that audio processing parameters are obtained based on the audio input signal by using a neural network whose coefficients are determined by reinforcement learning or are continuously adjusted.

[0097] The method 600 can be optionally supplemented with all the features, functions, and details described herein, even if such features, functions, and details are described in relation to the apparatus. The method can be supplemented with these features, functions, and details, both alone and in any combination.

[0098] Further embodiments In the following, several aspects of the invention are described which may be applied individually or in combination in the embodiments.

[0099] Context-dependent control parameters that are user adjustable or user parameters adjusted by the user can be integrated, for example, by analysis of input and output audio signals as shown in Figure 3 of sound processing during training of a self-reinforcement learning algorithm.

[0100] The input audio signal may include a sound environment, which allows changes in the sound environment and user adjustments to be inherently learned, for example, at run time.

[0101] From these data, a self-reinforcement learning algorithm can, for example, develop a set of parameters for each user for sound processing that applies the same control parameters in an automated manner under the same external conditions, but also allow further user adjustments in the context itself, which are integrated into the learning system (e.g., based on the principles of reinforcement learning). Thus, for example, machine learning systems and applications can be adapted to the user's sound preferences in a continuous learning process. For sound adaptation, the algorithm can be integrated and controlled, for example, when used in a hearing aid. Examples include multiband compression, interference noise reduction, or blind source separation with rate R, hearing threshold adjustment T, and band-dependent amplification G.

[0102] The input audio signal, sound processing parameters, and / or audio signals processed with the sound processing parameters can be stored, for example, in a cloud (e.g., a central data storage) to train a user profile. At the same time, sound processing parameters or user parameters selected by the user can be applied to the input audio signal. For example, the number of input parameters for reinforcement learning of a convolutional neural network (CNN) can be combined from, for example, multi-channel audio input (e.g., C = 4 channels) and audio output (e.g., C = 2 channels). The number of output parameters in the learned parameter set M can be combined from, for example, M = C * (T + R + G) + 2 time constants, where the number of values of G, R, and T can correspond to, for example, the number B of processing bands (e.g., B = 8).

[0103] In the following, several aspects of the invention are described which may be applied individually or in combination in the embodiments.

[0104] A possible implementation of the method, e.g., an apparatus in the field of sound control, is that a user carries a sound reproduction device (e.g., a hearable or earphone with additional functions) that includes, e.g., an integrated sound amplification and audio analysis system, as shown in FIG. 3 or FIG. 4. The sound amplification parameters can be controlled by the user, e.g., by an app (or application software), e.g., by using the user controls described above. In the background, audio analysis can continuously record an auditory scene, e.g., by microphonics, and analyze and evaluate, e.g., dynamics and / or frequency and / or spectral characteristics (e.g., audio analysis). In a specific auditory scene, e.g., while driving a car on a highway, the user can perform sound adjustments through the app and accordingly change the sound amplification parameters (e.g., parameters 272). A system (e.g., reinforcement learning 250) can establish an algorithmic connection between the user's parameter changes and the analysis of the auditory scene, from which artificial intelligence (AI) can develop a predictive model (e.g., represented by coefficients 256) that integrates further ad hoc user sound adjustments. This means that personalized AI control (AI, Artificial Intelligence) is enabled or provided by the device.

[0105] For example, if the user is again in the same auditory scene (in this case, a car traveling on a highway) at another time, the predictive model is applied and the sound amplification parameters (e.g., parameters 262) are implemented or provided in an automated manner by the system (e.g., by neural network 260 defined by coefficients 256). If the user again makes sound adjustments (e.g., via interface 230), this can be integrated, for example, in an ad-hoc manner into a self-learning system.

[0106] Below we describe several aspects of the present invention that may be applied individually or in combination in embodiments and that represent differences to the Github publication "liketohear-ai-pt."

[0107] According to an (optional) aspect of the invention, the predictive model is based on a real-time multi-dimensional optimization that analyzes all existing frequency bands.

[0108] According to an optional aspect of the invention, for example, reinforcement learning and unsupervised learning methods are used.

[0109] According to an (optional) aspect of the present invention, for example, the adjustment (or adjustments) of the processing parameter determination rules and / or audio processing parameters can be made continuously at run-time.

[0110] Below we describe several aspects of the invention that can be applied individually or in combination, for example in embodiments exhibiting differences relative to US Patent Application Publication No. 2015 / 195641.

[0111] Embodiments in accordance with the present invention, for example, are primarily concerned with intuitive and ergonomic user control of sounds in everyday acoustic environments and opt for the generalized adjustment option for the following reasons.

[0112] The division of a signal into individual "sound types" in real time is difficult to achieve in everyday acoustic situations. Therefore, the present invention does not apply this method, but covers multiple sound options with a two-dimensional parameter space.

[0113] User adjustments must be made separately for each object and each contextual situation in the signal separation. In everyday acoustic environments where hearing conditions change rapidly, user control becomes too complicated and therefore ergonomically inapplicable. The present invention allows users to perform complex sound adjustments (e.g., in means 230) through a simple, intuitive interface, such as a smartphone's 2D touch interface.

[0114] The sound characteristics of individual sounds may be perceived in combinations different from the user's preference, e.g., sounds such as music as foreground or background noise. Thus, in the present invention, the complexity of the auditory scene is adapted to the user's optimized perception of all existing sounds.

[0115] The adjustment of individual signals is not dynamically adapted to changing environmental conditions. So, for example, if you only play softly spoken language or soft music, a slight increase in the volume of the background noise may mean that the speech becomes unintelligible or the music is no longer audible.

[0116] Below we describe several aspects of the invention that can be applied individually or in combination, for example in embodiments exhibiting differences relative to U.S. Patent Application Publication No. 2020 / 0066264.

[0117] In US Patent Application Publication No. 2020 / 0066264, a processor controls the sound processing of the hearing aid due to the user's preferences and interests and historical activity patterns.

[0118] On the other hand, in an embodiment of the present invention, the sound processing of the hearing aid is controlled based on external conditions measured, for example in real time, as shown in FIG.

[0119] In summary, it must be stated that according to one aspect of the present invention, the above criteria or requirements are integrated into a learning method or device that learns in real time from user settings and applies them in an automated way in order to improve the individual perception of sounds or tones in the form of audio signals for the user. The present invention makes it possible to achieve a signal or sound reproduction optimized to the user's preferences.

[0120] Thus, according to one aspect of the present invention, the individual perception of sound and therefore the individual requirements for sound or euphony for the adjustment of a sound reproduction device can be considered to differ, inter alia, according to the following criteria:

[0121] ·Individuality · Situational needs External conditions According to one aspect of the present invention, embodiments in accordance with the present invention can take into account that sound perception varies from person to person.

[0122] For example, a conversation in a room with many people and a loud background noise is more difficult for some people than for others. Additionally, the same adjustments to sound reproduction, if required, are perceived differently.

[0123] According to one aspect of the invention, embodiments according to the invention can take into account that environmental parameters such as the auditory environment also have a significant influence on the control values for sound adjustment of the sound reproduction device.

[0124] In summary, embodiments according to the present invention can further be stated to provide an apparatus and method for performing sound processing based on immediately recorded or measured ambient noise. Based on these recordings and user parameters adjusted by the user, for example, a learning algorithm generates a predictive model that allows further adjustments in the context itself, which can be integrated into the learning system to improve the individual perception of sounds or tones in the form of audio signals for the user.

[0125] Although some aspects have been described in the context of an apparatus, it will be apparent that these aspects also represent descriptions of corresponding methods, and that blocks or devices of the apparatus also correspond to respective method steps or features of method steps. Similarly, aspects described in the context of a method step also represent descriptions of corresponding blocks or details or features of the corresponding apparatus.

[0126] Depending on specific implementation requirements, embodiments of the present invention can be implemented in hardware or software. Implementation can be performed using digital storage media such as floppy disks, DVDs, Blu-ray disks, CDs, ROMs, PROMs, EPROMs, EEPROMs, or flash memories, hard drives, or other magnetic or optical memories that store electronically readable control signals and that cooperate or can cooperate with a programmable computer system to perform the respective methods. Thus, the digital storage media can be computer-readable. Some embodiments according to the present invention include a data carrier having electronically readable control signals that can cooperate with a programmable computer system to perform one of the methods described herein.

[0127] Generally, embodiments of the present invention can be realized as a computer program product having program code operable to perform one of the above methods when the computer program product is run on a computer. The program code can, for example, be stored on a machine-readable carrier.

[0128] Other embodiments comprise the computer program for performing one of the methods described herein, wherein the computer program is stored on a machine readable carrier.

[0129] In other words, therefore, one embodiment of the inventive method is a computer program comprising a program code for performing one of the methods described herein when the computer program runs on a computer.A further embodiment of the inventive method is therefore a data carrier (or a digital storage medium or a computer readable medium) having recorded thereon a computer program for performing one of the methods described herein.

[0130] A further embodiment of the inventive method is, therefore, a data stream or a sequence of signals representing the computer program for performing one of the methods described herein, The data stream or the sequence of signals can be adapted to be transmitted via a data communication connection, for example the Internet.

[0131] A further embodiment comprises a processing means, for example a computer, or a programmable logic device, configured to or adapted to perform one of the methods described herein.

[0132] A further embodiment comprises a computer having installed thereon the computer program for performing one of the methods described herein.

[0133] Further embodiments according to the invention include an apparatus or system configured to transmit a computer program for performing at least one of the methods described herein to a receiver. The transmission may be, for example, electronic or optical. The receiver may be, for example, a computer, a mobile device, a memory device, or a similar device. The apparatus or system may, for example, include a file server for transmitting the computer program to the receiver.

[0134] In some embodiments, a programmable logic device (e.g., a field programmable gate array, FPGA) can be used to perform some or all of the functions of the methods described herein. In some embodiments, a field programmable gate array can cooperate with a microprocessor to perform one of the methods described herein. In general, the methods are preferably performed by any hardware apparatus. This may be universally applicable hardware, such as a computer processor (CPU), or hardware specific to the method, such as an ASIC.

[0135] The above-described embodiments are merely illustrative of the principles of the present invention. It should be understood that modifications and variations of the arrangements and details described herein will be apparent to those skilled in the art. Therefore, the present invention is limited only by the scope of the appended claims, and not by the specific details presented by the description and interpretation of the embodiments herein.

Claims

1. 1. An apparatus (100; 200; 300; 400; 500) for determining audio processing parameters (120; 262; 362; 462; 562) in response to at least one audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516), comprising: the device (100; 200; 300; 400; 500) is configured to determine at least one coefficient (142; 256; 356; 456; 556) of a processing parameter determination rule (140; 250; 350; 450; 550) in a user-specific manner based on an audio signal (217, 218, 219; 313, 317, 318, 319, 342; 417; 517) acquired during user operation, the device (100; 200; 300; 400; 500) is configured to obtain the audio processing parameters (120; 262; 362; 462; 562) by using the processing parameter determination rule (140; 250; 350; 450; 550) based on the audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516); the device is configured to determine, in response to the at least one audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516), a database (252; 352; 452; 552) such that entries in the database (252; 352; 452; 552) represent the audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516); said device being configured to determine said database (252; 352; 452; 552) in response to an audio output signal (218, 219, 313, 318, 319, 342) obtained in response to user parameters, such that entries in said database (252; 352; 452; 552) represent said audio output signal (218, 219, 313, 318, 319, 342); The apparatus (100; 200; 300; 400; 500) is configured to adjust the at least one coefficient (142; 256; 356; 456; 556) of the processing parameter determination rule (140; 250; 350; 450; 550) based on the database (252; 352; 452; 552) obtained by the apparatus, in order to adjust the processing parameter determination rule (140; 250; 350; 450; 550) in a user-specific manner to obtain audio processing parameters (120; 262; 362; 462; 562) adjusted in a user-specific manner.

2. The device (100; 200; 300; 400; 500) of claim 1, configured to determine the database (252; 352; 452; 552) in response to user parameters (232; 332; 432, 433; 532) adjusted by a user such that entries in the database (252; 352; 452; 552) represent the user parameters (232; 332; 432, 433; 532) adjusted by the user.

3. 3. The apparatus (100; 200; 300; 400; 500) of claim 2, wherein the apparatus is configured to determine the database (252; 352; 452; 552) such that the database (252; 352; 452; 552) represents an allocation between different audio input signals (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516) and respective user parameters (232; 332; 432, 433; 532) adjusted by the user.

4. The apparatus (100; 200; 300; 400; 500) according to any one of claims 1 to 3, wherein the apparatus is configured to determine the database (252; 352; 452; 552) such that the database (252; 352; 452; 552) represents an allocation between different audio output signals (218, 219, 313, 318, 319, 342) and respective user parameters (232; 332; 432, 433; 532) adjusted by the user.

5. The apparatus (100; 200; 300; 400; 500) according to any one of claims 1 to 4, wherein the apparatus is configured to provide or adjust the processing parameter determination rules (140; 250; 350; 450; 550) based on the database (252; 352; 452; 552).

6. The apparatus (100; 200; 300; 400; 500) according to any one of claims 1 to 5, configured to determine and / or adjust the at least one coefficient (142; 256; 356; 456; 556) of the processing parameter determination rule (140; 250; 350; 450; 550) based on at least one audio processing parameter (272; 372; 473; 573) corrected and / or compensated by a user.

7. The device (100; 200; 300; 400; 500) according to any one of claims 1 to 6, wherein the device is configured to perform audio processing (220; 320; 420; 520) based on the audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516) and the audio processing parameters (120; 262; 362; 462; 562) to obtain the audio signal (217, 218, 219; 313, 317, 318, 319, 342) adjusted in a user-specific manner.

8. 8. The apparatus (100; 200; 300; 400; 500) of claim 7, wherein the apparatus is configured to determine the coefficients (142; 256; 356; 456; 556) of the processing parameter determination rule (140; 250; 350; 450; 550) by using a comparison of the audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516) with an audio output signal (218, 219, 313, 318, 319, 342) provided by the audio processing (220; 320; 420; 520) by using the audio processing parameters (120; 262; 362; 462; 562).

9. The device (100; 200; 300; 400; 500) according to any one of claims 1 to 8, wherein the device is configured to provide the user parameters (232; 332; 432, 433; 532) adjusted by the user as output quantities instead of the audio processing parameters (120; 262; 362; 462; 562), and the user parameters (232; 332; 432, 433; 532) adjusted by the user comprise volume parameters and / or sound parameters and / or equalizer parameters.

10. The device (100; 200; 300; 400; 500) according to claim 7 or 8, configured to combine the user parameters (232; 332; 432, 433; 532) with the audio processing parameters (120; 262; 362; 462; 562) to obtain composite parameters (272; 372; 472, 473, 474; 572, 573) of the audio processing (220; 320; 420; 520), which are provided as output quantities.

11. The apparatus (100; 200; 300; 400; 500) according to any one of claims 1 to 10, wherein the apparatus is configured to perform an audio analysis of the audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516) to provide an audio input signal analysis result for determining the at least one coefficient (142; 256; 356; 456; 556) of a processing parameter determination rule (140; 250; 350; 450; 550).

12. The apparatus (100; 200; 300; 400; 500) according to any one of claims 1 to 11, wherein the apparatus is configured to perform an audio analysis of the audio output signal (342) and provide an audio output signal analysis result for determining the at least one coefficient (142; 256; 356; 456; 556) of the processing parameter determination rule (140; 250; 350; 450; 550).

13. 13. The apparatus (100; 200; 300; 400; 500) according to any one of claims 1 to 12, wherein the audio processing parameters (120; 262; 362; 462; 562) comprise at least one multiband compression parameter R and / or at least one hearing threshold adjustment parameter T and / or at least one band dependent amplification parameter G and / or at least one interference noise reduction parameter and / or at least one blind source separation parameter and / or at least one sound direction parameter and / or at least one binaural parameter and / or at least one parameter of an adaptive filter.

14. The apparatus (100; 200; 300; 400; 500) according to any one of claims 1 to 13, comprising a neural network (260; 360; 460; 560) configured to obtain the audio processing parameters (120; 262; 362; 462; 562) by using the processing parameter decision rules (140; 250; 350; 450; 550).

15. The apparatus (100; 200; 300; 400; 500) according to any one of claims 1 to 14, configured to provide and / or adjust the processing parameter determination rules (140; 250; 350; 450; 550) based on a multidimensional parameter space determined by reinforcement learning methods and / or unsupervised learning methods and / or multivariate prediction methods and / or multivariate regression to determine the audio processing parameters (120; 262; 362; 462; 562).

16. The device (100; 200; 300; 400; 500) according to any one of claims 1 to 15, wherein the device is configured to obtain the user parameters (232; 332; 432, 433; 532) adjusted by the user from an interface.

17. The apparatus (100; 200; 300; 400; 500) according to any one of the preceding claims, wherein the audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516) comprises a multi-channel audio signal or at least two audio channels.

18. The apparatus (100; 200; 300; 400; 500) according to any one of claims 1 to 17, wherein the apparatus is configured to perform audio processing (220; 320; 420; 520) separately for at least four frequency bands of the audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516).

19. The apparatus (100; 200; 300; 400; 500) according to any one of claims 1 to 18, wherein the apparatus is configured to determine the at least one coefficient (142; 256; 356; 456; 556) of the processing parameter determination rule (140; 250; 350; 450; 550) in a user-specific manner during user operation in order to obtain the audio processing parameters (120; 262; 362; 462; 562) in real time.

20. The apparatus (100; 200; 300; 400; 500) of claim 6, wherein the apparatus is configured to determine the at least one coefficient (142; 256; 356; 456; 556) of the processing parameter determination rule (140; 250; 350; 450; 550) in a user-specific manner during user operation in order to determine and / or adjust the corrected audio processing parameter (120; 262; 362; 462; 562) in real time.

21. Includes audio processing comprising an apparatus for determining audio processing parameters according to any one of claims 1 to 20, The hearing aid, wherein the audio processing is configured to process an audio input signal in response to the audio processing parameters.

22. A method (600) for determining audio processing parameters in response to at least one audio input signal, comprising: determining, in a per-user manner, at least one coefficient of a processing parameter determination rule based on an audio signal acquired during user operation; obtaining audio processing parameters by using the processing parameter determination rules based on the audio input signal; Including, a database (252; 352; 452; 552) is determined in response to said at least one audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516) such that entries in said database (252; 352; 452; 552) represent said audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516); said database (252; 352; 452; 552) is determined in response to an audio output signal (218, 219, 313, 318, 319, 342) obtained in response to user parameters, such that entries of said database (252; 352; 452; 552) represent said audio output signal (218, 219, 313, 318, 319, 342); a processing parameter determination rule for adjusting the audio processing parameters of the audio signal from the audio signal to the audio signal;

23. 23. A computer program having a program code for performing the method according to claim 22 when the computer program runs on a computer.

Citation Information

Patent Citations

  • Learning hearing aid

    JP2015130659A

  • Deep learning-based sound quality characteristic processing method and system

    JP2021525493A

  • Hearing aid for recording data and learning therefrom

    US20060222194A1

  • User adjustment interface using remote computing resource

    US20190149929A1

  • Methods and Systems for Automatically Equalizing Audio Output based on Room Position

    US20200220511A1