Apparatus and method for determining audio processing parameters

DE502022005217D1Active Publication Date: 2025-09-18FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
DE502022005217
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2021-05-17
Filing Date
2022-05-16
Publication Date
2025-09-18
Estimated Expiration
2042-05-16

AI Technical Summary

Technical Problem

Conventional sound reproduction devices fail to adapt audio processing parameters to individual user preferences and environmental changes, requiring repetitive manual adjustments and often providing inadequate sound quality due to generalized settings.

Method used

A device and method that determines audio processing parameters in real-time using user-specific settings based on audio input signals, employing a processing parameter determination rule adjusted by user operation and reinforcement learning to adapt to individual habits and preferences.

Benefits of technology

Provides user-specific audio processing parameters that dynamically adjust to varying listening environments and preferences, enhancing audio quality and reducing the need for manual adjustments.

✦ Generated by Eureka AI based on patent content.
Patent Text Reader
Need to check novelty before this filing date? Find Prior Art

Description

Technical area

[0001] Embodiments according to the present invention relate to an apparatus and a method for determining audio processing parameters depending on at least one audio input signal.

[0002] Embodiments according to the invention relate to a device and a method with artificial intelligence, for example in a sound reproduction device, which can analyze audio signals and assign or combine them with user-specific settings during user operation.

[0003] Embodiments further relate to concepts for determining audio processing parameters based on audio signals obtained during user operation. Background of the invention

[0004] The individual perception of sound and thus the individual requirements for the sound or euphony in their adaptation of sound reproduction devices differ according to the following criteria, among others: Individuality Situational needs External conditions

[0005] Sound perception differs from person to person. For example, a conversation with one person in a crowded room may be more difficult for one person than for another. Likewise, the same sound reproduction setting is perceived differently depending on individual needs. Environmental parameters, such as the auditory environment, also significantly influence the control values ​​for sound adjustment of a sound reproduction device.

[0006] Current sound reproduction devices offer specific sound adjustments that are not applied automatically. Sound reproduction devices, such as portable hearing aids such as headphones, headsets, or hearing aids, often only offer volume control and an equalizer for sound adjustment. Sound adjustments, such as increasing the volume or adjusting the higher or lower frequencies, are performed once by the user. It has been recognized that to achieve consistently good audio quality, these settings must be repeated for each subsequent sound reproduction.

[0007] It was recognized that with conventional concepts, not only does the process of sound adaptation have to be repeated for different sound reproductions, but even with sound reproduction devices, changes in the auditory environment are not adaptively adjusted, for example, to ambient noise. It was recognized that even a relatively minor change in background noise can increase the listening effort required for speech understanding.

[0008] Furthermore, it was recognized that with conventional concepts, sound adjustments can only be made based on the sound presets specified by the manufacturer. It was discovered that these do not always meet individual needs. For example, there are settings such as "Music," which do not take into account the preferred musical taste and personal intention when listening to music. For example, the expectations regarding the sound experience of opera singing differ fundamentally from those of techno. However, the presets in the "Music" listening program are based on generalized assumptions by the manufacturer that may not meet the expectations regarding the sound experience of either opera singing or techno, thus providing the user with inadequate sound reproduction.

[0009] Modern sound reproduction devices for assistive listening, such as hearing aids, can cost several thousand euros depending on the features, so expectations for the device are correspondingly high. Hearing aid fittings are generally performed under laboratory conditions, usually with only two speakers and a very limited range of sounds, such as pure tones, noise, and speech. Complex noise situations, such as those at intersections, cannot be simulated in a hearing laboratory, thus leading to frustration for hearing aid wearers and unsatisfactory results in everyday use.

[0010] In learning applications for sound reproduction, such as the Github publication "liketohear-ai-pt," situation-dependent parameter changes of a hearing aid algorithm recorded by users in a file and the recorded frequency spectrum analysis corresponding to the situation are processed with a self-learning algorithm. The algorithm establishes the relevance of a specific frequency spectrum for the user's decision and automatically selects the corresponding parameters as the basis for a prediction model. In a second step, the prediction model is applied to the previously recorded frequency spectrum analysis. It was recognized that this learning application for sound reproduction cannot represent the complexity of the frequency spectrum, so further user adjustments are repeatedly necessary.

[0011] US 2019 / 0149929 A1 describes a user settings interface using a remote computing resource. A system includes a mobile device communicating with a hearing assistance device or a remote server. The mobile device interprets an acoustic environment and sends information about the environment to a remote server. The remote server determines information and sends it to the mobile device for use in a user interface. The mobile device receives a user selection of hearing assistance parameter information to be sent to the hearing assistance device.

[0012] In view of the above, there is a need for a concept for determining audio processing parameters at runtime that provides an improved trade-off between usability, achievable audio quality, and implementation effort. Summary of the invention

[0013] This object is achieved by the subject matter of the independent patent claims. A core idea of ​​embodiments of the present invention is to recognize that sound adjustments intuitively performed by users can be implemented in runtime and integrated into the learning system in real time.

[0014] An embodiment according to the present invention comprises a device for determining audio processing parameters, for example parameters for audio processing, as a function of at least one audio input signal, for example coming from an audio input, wherein the device is designed to determine at least one coefficient of a processing parameter determination rule in a user-specific manner based on audio signals obtained during user operation, and wherein the device is designed to obtain the audio processing parameters using the processing parameter determination rule based on the audio input signal.Coefficients of a processing parameter determination rule can, for example, be coefficients of a neural network that receives the audio input signal, or input signal parameters extracted therefrom, as input variable and provides the audio processing parameters as output variable. In other words, the coefficients of the processing parameter determination rule can, for example, be determined on a user-specific basis based on input audio signals received during user operation, for example during user operation. Furthermore, the device can be designed to obtain the audio processing parameters, for example using the processing parameter determination rule defined by the at least one coefficient, based on the audio input signal.

[0015] This embodiment is based on the core idea that a user-specific setting of one or more coefficients of the processing parameter determination rule based on audio signals obtained during user operation makes it possible to adapt the processing parameter determination rule to the user's individual habits and preferences. By using audio signals obtained during user operation for the user-specific setting of the coefficients of the processing parameter determination rule, it is possible to ensure that the coefficients are well adapted to the (concrete) listening situations in which the user typically finds himself.Thus, for example, it is no longer necessary to preclassify an acoustic environment (for example, into a general category of "music" and a general category of "speech"). Instead, the coefficients can be adapted to the actual listening environments in which the user listens to music or speech, for example, and also to the user's individual needs. For example, by appropriately selecting the coefficients of the processing parameter determination rule, a direct and user-specific determination of audio processing parameters can be achieved. For example, the processing parameter determination rule adapted by coefficients requires a direct determination of the audio processing parameters without categorizing the acoustic environment into one of several statically predetermined categories.Rather, coefficients of the processing parameter determination rule can be adjusted based on the audio signals obtained during user operation, so that the listening environments relevant to the user, in which the user desires different audio processing parameters, can be distinguished as "hard" or "soft" (e.g., with a smooth transition).

[0016] Thus, by taking into account the audio signals received during user operation (and by appropriately adjusting the coefficients of the processing parameter determination rule), the inventive concept enables, for example, the provision of completely different audio processing parameters in the presence of speech in different acoustic environments where the user is located (e.g., a noisy open-plan office, a private office, a street intersection with numerous trucks, a street intersection with tram traffic, etc.). The provided parameters are then typically oriented towards the settings desired by the user in the respective situations.

[0017] In this respect, the invention concept provides, with reasonable effort, audio processing parameters that are adapted to the real life of an individual user and his or her specific preferences.

[0018] According to a further embodiment, the device is configured to determine a database depending on user parameters set by the user, so that entries in the database describe the user parameters set by the user. For example, the database can be created in real time during user operation, and a predictive model can be determined. Furthermore, the database can be used to determine the coefficients of the processing parameter determination rule by using the database to contain information about the user parameters. For example, the database can also contain personalized control settings that can be linked to the user parameters.The user parameters set by the user can, for example, replace the audio processing parameters as the output variable, or they can modify the audio processing parameters so that the database entries represent, for example, the user parameters set by the user. For example, the database is at least partially integrated into reinforcement learning, which, for example, uses the user parameters set by the user.

[0019] By creating a database whose entries describe the user parameters set by the user, the coefficients of the processing parameter determination rule can, for example, be successively improved or optimized. The user parameters set by the user (typically in different acoustic environments), which form the database and are stored, for example, in a database or other storage structure, can represent target values ​​of audio processing parameters. If, for example, user parameters are assigned to audio signals (or audio signal properties) of the respective acoustic environment in which the user selected the user parameters, this database can be used to determine the coefficients of the processing parameter determination rule.By defining a database that grows larger with the duration of user use, for example, it is possible to achieve an increasingly larger database for the (automatic) determination (or improvement) of the coefficients of the processing parameter determination rule over time, which enables increasing refinement or improvement of these coefficients (e.g., based on an increasingly larger database of different listening environments in which the user has been). Thus, the user experience can be continually improved by creating and continuously expanding the database.

[0020] According to the invention, the device is designed to determine a database as a function of the at least one audio input signal, such that entries in the database represent or describe the audio input signal. For example, the database can be used to determine the coefficients of the processing parameter determination rule. In other words, for example, personalized control settings, such as the user parameters set by the user, are initially stored, which are then expanded with sound information from the auditory environment as an external framework. This makes it possible to create a database that, for example, uses reinforcement learning to provide coefficients for the processing parameter determination rule.

[0021] According to a further embodiment, the device is configured to determine the database such that the database describes an association between various audio input signals and respective user parameters set by the user. In other words, the device can, for example, associate the external framework conditions based on the audio input signal with the personalized control settings, for example, the user parameters set by the user. This means that the association can, for example, serve as a basis for the prediction model, which can be modified, for example, ad hoc, by further sound adjustments by the user, for example by integrating the respective user parameters set by the user with the database (and then, for example, redetermining or improving the coefficients of the processing parameter determination rule).For example, the auditory scene can be continuously recorded, analyzed, and / or evaluated in the background via the audio input using microphones, generating, for example, an analysis of the auditory scene based on its dynamics, frequency, and / or spectral properties. The analysis result of the auditory scene can, for example, be integrated into the database as an environmental parameter and assigned to the user parameter to obtain a link between the user parameter and the audio input signal in the auditory environment for that particular time.

[0022] According to the invention, the device is designed to determine a database, for example for determining the coefficients of the processing parameter determination rule, as a function of an audio output signal, so that entries in the database describe or represent the audio output signal. By determining the database as a function of at least one audio input signal and one audio output signal, the processing parameter determination rule, for example of reinforcement learning, can use the database to determine coefficients of the processing parameter determination rule, for example for a neural network. The coefficients of the processing parameter processing rule can be obtained, for example, by jointly processing an audio input signal and an associated output signal or by comparing the audio output signal with the audio input signal.

[0023] According to a further embodiment, the device is configured to determine the database such that the database describes a mapping between different audio output signals and respective user parameters set by the user. In other words, the database describes a mapping between different audio input signals, between different audio output signals and respective user parameters set by the user, in order to be able to determine coefficients of the processing parameter determination rule. Using the created database, sound processing can be integrated into the training of a self-reinforced learning algorithm, for example by analyzing the incoming and outgoing audio signals. For example, the incoming audio signal or the audio input signal can contain the sound environment, for example the auditory environment.In other words, using the created database, for example by analyzing the incoming and outgoing audio signal, the coefficients of the processing parameter determination rule can be selected such that the desired relationship between the audio input signal and the audio output signal is at least approximately obtained by the processing parameter determination rule.

[0024] According to the invention, the device is designed to adapt the at least one coefficient of the processing parameter determination rule based on the database acquired by the device in order to adapt the processing parameter determination rule to the individual user in order to obtain user-specifically adapted audio processing parameters. In other words, for example, the reinforcement learning user model is adapted based on artificial intelligence in order to obtain user-specifically adapted audio processing parameters or a user-specifically adapted audio signal. For example, it is thus possible to inherently learn and adapt changes in the sound environment, for example the auditory environment, and the user settings, for example the user parameters, at runtime.For example, user-specific audio processing parameters can enable the production of user-specific audio signals during user operation when processing the audio input signal using the audio processing parameters. In other words, a user-specific set of parameters for sound processing can be obtained or developed from the database. This set automatically applies the same control parameters under the same external conditions, but also allows for further user adjustments in the situation itself, which are integrated into the device as a learning system. For example, the learning system and the application can adapt to the user's sound preferences in a continuous learning process.

[0025] According to a further embodiment, the device is configured to provide and / or adapt the processing parameter determination rule based on the database. For example, the device can use the database, for example using reinforcement learning, to provide the processing parameter determination rule in order to obtain user-specifically adapted audio signals using the audio processing parameters, for example, during user operation.

[0026] According to a further embodiment, the device is configured to determine and / or adapt the at least one coefficient of the processing parameter determination rule based on at least one audio processing parameter corrected and / or modified by a user. As already mentioned, the device can be configured to consider or adjust user adjustments of the user parameters during user operation and, for example, to allow further user adjustments of the user parameters at a later time and at the same location or in the same sound environment, so that the previous user parameters are adjusted and / or overwritten with newly adjusted user parameters.In other words, coefficients of the processing parameter determination rule can be corrected by a user and / or, for example, modified audio processing parameters can be determined, for example depending on the sound environment at the respective time in which the user is located.

[0027] According to a further embodiment, the device is configured to perform audio processing, for example, a parameterized audio processing rule, based on the audio input signal and based on the audio processing parameters in order to obtain the user-customized audio signals, for example, taking into account user modifications of the audio processing parameters. In other words, the device can provide a user-customized audio signal for the audio output by means of optional audio processing of the audio input signal and the audio processing parameters. Thus, for example, the audio processing can be integrated into the device, thereby obtaining an efficient system. The audio processing can optionally also be included in the determination of the audio processing parameters.

[0028] According to a further embodiment, the device is configured to determine the coefficients of the processing parameter determination rule using a comparison of the audio input signal and an audio input signal supplied using the audio processing parameter, for example, taking into account user modifications of the audio processing parameters. In other words, the determination of the coefficients of the processing parameter determination rule can be based on a comparison between the audio input signal and the direct audio output signal or the audio output signal supplied by the audio processing. For example, an audio analysis of the audio input signal or an audio analysis of the audio output signal can optionally be performed before or after using the comparison in order to determine the coefficients of the comparison parameter determination rule based on an audio analysis result of the audio signals.Determining the coefficients of the parameter determination rule using such a comparison yields particularly reliable and robust results, since the actual audio signal output to the user can be used as the criterion for determining the coefficients of the parameter determination rule. The criterion that the audio output signal should correspond to what the user desires is more meaningful and robust than simply optimizing the audio processing parameters themselves.

[0029] According to a further embodiment, the device is configured to provide the user parameters set by the user as an output variable instead of the audio processing parameters, wherein the user parameters set by the user include volume parameters and / or tone parameters and / or equalizer parameters. In other words, user parameters can include, for example, filter parameters for sound shaping and / or for equalizing audio frequencies. By providing the user parameters set by the user as an output variable, short-term user intervention is enabled, for example, resulting in a particularly good user experience. User intervention can then be used additionally to improve the coefficients in order to avoid future user interventions wherever possible (and instead automatically achieve a setting adapted to the user's wishes).

[0030] According to a further embodiment, the device is designed to combine the user parameters with the audio processing parameters, for example, by addition, to thereby obtain combined audio processing parameters and provide them as an output variable. Combined parameters can, for example, comprise user parameters and audio processing parameters that are provided in combination to the audio processing or are combined using the audio processing and provided as an output variable, for example, to reinforcement learning. Accordingly, rapid user intervention is possible, and the audio processing can thus be adapted to the user's preferences.

[0031] According to a further embodiment, the device is configured to perform an audio analysis of the audio input signal in order to provide an audio input signal analysis result for determining the at least one coefficient of a processing parameter determination rule, for example, using the processing parameter determination rule. For example, the processing parameter determination rule can define a derivation rule for deriving the audio processing parameters from the audio input signal analysis result. The audio analysis of the audio input signal can provide audio input signal analysis results, for example, in the form of information about spectral properties and / or dynamics and / or frequency of the audio input signal, or also information about intensity values ​​per band.The audio input signal analysis results can, for example, be provided as input variables for determining one or more coefficients of the processing parameter determination rule, for example using reinforcement learning. Embodiments further provide that the audio analysis analyzes and evaluates the audio input signal coming from the audio input in advance in order to provide it to the processing parameter determination rule, although this is not mandatory. For example, it is possible to obtain additional information about spectral properties of the audio input signal as an audio input signal analysis result. Furthermore, by using an audio input signal analysis result, the processing parameter determination rule can be designed more simply than if, for example, the entire audio input signal were used to determine audio processing parameters.For example, parameters or values ​​of the audio input signal analysis result can efficiently describe the essential characteristics of the audio input signal, so that the processing parameter determination rule has a comparatively small number of input variables (namely, for example, the parameters or values ​​of the audio input signal analysis result) and is therefore comparatively easy to implement. Thus, good results can be achieved with little effort.

[0032] According to a further embodiment, the device is configured to perform an audio analysis of the audio output signal in order to provide an audio output signal analysis result, for example in the form of information about spectral properties of the audio input signal, for determining the at least one coefficient of the processing parameter determination rule, for example using the processing parameter determination rule. In other words, the device is configured to perform an audio analysis before the processing parameter determination rule or after the processing parameter determination rule in order to provide either an audio input signal analysis result or an audio output signal analysis result, or both, for determining the coefficient of the processing parameter determination rule.For example, by determining the audio output signal analysis result, it is particularly easy to compare the audio input signal and the audio output signal, whereby, for example, values ​​or parameters of the audio output signal analysis result can describe the characteristic properties of the audio output signal particularly efficiently (or in a particularly compact form). Thus, a determination or optimization of the coefficients of the processing parameter determination rule is particularly efficient, whereby the processing desired by the user can be achieved, for example, efficiently by evaluating the audio output signal analysis result, or whereby a comparison between the audio input signal analysis result and the audio output signal analysis result can allow a conclusion to be drawn about the coefficients of the processing parameter determination rule.

[0033] According to a further embodiment, the audio processing parameter(s) comprise(s) at least one multiband compression parameter R, and / or at least one hearing threshold adaptation parameter T, and / or at least one band-dependent gain parameter G, and / or at least one noise reduction parameter and / or at least one blind source separation parameter. Furthermore, the audio processing parameters may comprise at least one sound direction parameter, and / or binaural parameters, and / or parameters relating to the number of different speakers, and / or parameters of adaptive filters in general, for example, reverberation suppression, feedback, echo cancellation, or active noise cancellation (ANC).For example, using a sound direction parameter, the directivity of the sound source can be selected or adjusted so that sound is only processed from the desired direction, such as the person speaking in a conversation, for the combination of audio processing parameters. It has been recognized that such audio processing parameters can efficiently influence audio signal processing. Even with a small number of parameters, which can be easily determined using a processing parameter determination rule, it is possible to influence audio signal processing over a wide setting range.

[0034] According to a further embodiment, the device can comprise a neural network that, for example, implements the processing parameter determination rule so that the at least one coefficient is defined, or preferably a plurality of coefficients are defined that are designed to obtain the audio processing parameters using the processing parameter determination rule. Furthermore, the neural network can be designed to obtain the audio processing parameters based on the audio input signal directly from the audio input or by means of the intermediate audio analysis as an analyzed audio input signal. It has been recognized that a neural network is well suited for determining the audio processing parameters and can be easily adapted to the personal perception of the individual user through the coefficients.The neural network, whose edge weights can be defined, for example, by the coefficients of the processing parameter determination rule, can be adapted to the user's needs by selecting the coefficients (which can be done, for example, using a training rule). The coefficients can be successively improved, for example, as additional user settings are made. This allows for results that offer a very good user experience.

[0035] According to a further embodiment, the device is configured to provide and / or adapt the processing parameter determination rule based on a reinforcement learning method, and / or based on a reinforcement learning method, and / or based on an unsupervised learning method, and / or based on a multivariate prediction method, and / or based on a multidimensional parameter determined using multivariable regression, in order to determine the audio processing parameter. The processing parameter determination rule can, for example, provide coefficients for the neural network that are based, for example, on the reinforcement learning method. The multivariate prediction method can, for example, comprise a prediction of frequency bands and / or a prediction of input / output characteristics or input / output characteristics according to the user parameters.Furthermore, the method can, for example, analyze all available frequency bands using multivariable regression to define a multidimensional parameter space. A multidimensional parameter space can be understood, for example, as a two-dimensional parameter setting with a graphical interface in which the user parameters can be set and continuously adjusted by the user, for example, using sliders or a point on a coordinate system whose axes contain or are assigned to volume and tone settings.Using the methods listed above, the device can determine the audio processing parameters so that, for example, a learning algorithm sets user-specific audio processing parameters, or so that the audio processing parameters provided by applying the processing parameter determination rule approach the audio processing parameters corrected by the user with increasing learning progress, or so that the processing parameter determination rule adapts in a continuous learning process, for example, depending on user adjustments to the audio processing parameters. As expected, for example, access of the methods to the database or data storage is unrestricted (so that, for example, as the database size increases, increasingly better coefficients can be determined using the aforementioned learning methods).

[0036] According to a further embodiment, the device is designed to receive the user parameters set by the user, for example via or by means of an interface, for example from a user interface, an intuitive and / or ergonomic user control, such as a 2D space on a smartphone display. In other words, the device can comprise an interface (for example an electrical interface or a human-machine interface) to be able to set the user parameters. Preferably, a visual user control can comprise a volume adjustment, for example by means of a slider for louder and quieter and / or a treble and bass control. In this way, the adjustment of the parameters can be made very easy for the user, wherein it has been recognized that this simple sound adjustment already results in a good auditory impression in many cases.

[0037] According to a further embodiment, the audio input signal comprises a multi-channel audio signal, for example, with at least four channels or at least two audio channels. For example, the audio input signal can be provided by the audio input, for example, from, via, or by means of a microphone. Furthermore, the audio input signal can contain information such as the number of channels and / or the number of frequency bands. The use of multi-channel signals allows, for example, a localization of desired and / or interfering sound sources as well as a consideration of the directions of the desired or interfering sound sources when determining the audio processing parameters or the coefficients of the processing parameter determination rule.

[0038] According to a further embodiment, the device is designed to perform audio processing separately for at least four frequency bands of the audio input signal. This ensures that frequency selectivity is provided to enable analysis of each individual frequency, for example, if the audio input signal comprises a multi-channel audio signal. Taking into account the different intensities in different frequency bands makes it possible to efficiently accommodate different acoustic environments and also to efficiently consider the user's specific preferences regarding the frequency response.

[0039] According to a further embodiment, the device is designed to determine the at least one coefficient of the processing parameter determination rule on a user-specific basis, for example continuously, during user operation, for example in real time, in order to obtain the audio processing parameters in real time, for example at runtime during user operation, and / or to determine and / or adapt the changed audio processing parameters in real time. In other words, the device is designed, for example, to determine and / or adapt the audio processing parameters in real time, so that the device, as a learning system, carries out this learning process in real time, for example during user operation. In other words, in the present invention, for example, the sound processing is controlled based on external framework conditions measured in real time.Thus, an analysis of all available frequency bands is also carried out in real time, so that the prediction model can be provided in real time based on a multidimensional optimization, that is, for example, an optimization in which the audio processing parameters are determined based on the analyzed frequency bands and the user parameters stored in the data memory.

[0040] According to a further embodiment, the present invention comprises a hearing aid, wherein the hearing aid has audio processing and wherein the hearing aid has a device for determining audio processing parameters, wherein the audio processing is designed to process an audio input signal depending on the audio processing parameters. For example, the hearing aid can implement or integrate the device to improve the user's individual perception of sound or tones in the form of audio signals. It has been shown that the device described herein is particularly well suited for use in a hearing aid, and that the auditory impression can be significantly improved by using the inventive concept.

[0041] An embodiment according to the present invention comprises a method for determining audio processing parameters as a function of at least one audio input signal, wherein the method comprises a user-specific determination of at least one coefficient of a processing parameter determination rule based on audio signals obtained during user operation, and obtaining audio processing parameters using the processing parameter determination rule based on the audio input signal. The method is based on the same considerations as the device described above and can optionally be supplemented by all features, functionalities, and details that are also described herein with regard to the device according to the invention. The method can be supplemented by the aforementioned features, functionalities, and details both individually and in combination.

[0042] A further embodiment according to the present invention comprises a computer program with a program code for carrying out the method when the program is running on the computer. Short description of the characters

[0043] Exemplary embodiments are explained below with reference to the accompanying drawings. They show: Fig. 1 shows a schematic block diagram of a device that determines audio processing parameters as a function of at least one audio input signal; Fig. 2 shows a schematic block diagram of a device according to an embodiment that determines audio processing parameters as a function of at least one audio input signal and by means of reinforcement learning, based on an audio input signal and an audio output signal; Fig. 3 shows a schematic block diagram of a device according to an embodiment that determines audio processing parameters as a function of at least one audio input signal and by means of reinforcement learning, based on an audio analysis of the audio input signal and an audio analysis of the audio output signal; Fig.Fig. 4 shows a schematic block diagram of a device that determines audio processing parameters as a function of at least one audio input signal and by means of reinforcement learning, based on an audio analysis of the audio input signal and on user parameters set by the user; Fig. 5 shows a schematic block diagram of a device that determines audio processing parameters as a function of at least one audio input signal and by means of reinforcement learning, based on an audio input signal and on user parameters set by the user; and Fig. 6 shows a schematic flow diagram of a method for determining audio processing parameters. Detailed description of embodiments of the invention

[0044] Before exemplary embodiments of the present invention are explained in more detail below with reference to the drawings, it is pointed out that identical, functionally equivalent or equivalent elements, objects and / or structures in the different figures are provided with the same reference numerals, so that the description of these elements shown in different exemplary embodiments is interchangeable or can be applied to one another.

[0045] The embodiments described below are described in conjunction with numerous details. However, embodiments may also be implemented without these detailed features. Furthermore, for clarity, embodiments are described using block diagrams instead of detailed illustrations. Furthermore, details and / or features of individual embodiments may be readily combined with one another, unless explicitly described otherwise.

[0046] Fig. 1shows a schematic block diagram of a device 100 for determining audio processing parameters 120, which are represented on the output side of the device 100, as a function of at least one audio input signal 110, which is represented on the input side of the device 100. The exemplary schematic representation of the device 100 includes, for example, a determination of coefficients, which is represented by the coefficient determination block 130, so that coefficients 132 of the coefficient determination 130 can be provided to the processing parameter determination rule 140.For example, the audio input signal 110 can be used directly by the processing parameter determination rule 140 to obtain the coefficients 142 of the processing parameter determination rule 140, and / or used as an audio signal 112 obtained during user operation by the coefficient determination rule 130 to provide the coefficients 132 of the coefficient determination rule 130. For example, the coefficient determination rule 130 can be performed on a user-specific basis during user operation, so that the coefficients 132 of the coefficient determination rule 130 are provided to the processing parameter determination rule 140 to obtain the audio processing parameters 120 using the processing determination rule 140 based on the audio input signal 110.

[0047] Thus, the coefficients of the processing parameter determination rule can, for example, be set such that the processing parameter determination rule, based on the audio input signal and using the coefficients as output, provides audio processing parameters which, when used in audio processing, result in an audio output signal that meets user expectations.

[0048] Fig. 2 shows a schematic block diagram of a device 200 according to an embodiment. The illustrated device 200 for determining audio processing parameters comprises, for example, an audio input 210, an audio processor 220, a user controller 230, an audio output 240, a processing determination rule (or processing parameter determination device) in the form of reinforcement learning 250, and a neural network 260.

[0049] The audio input 210 may, for example, comprise a microphone or other audio capture device and may contain, for example, information about the number of channels, for example, "C," and / or information about the number of frequency bands, for example, "B." For example, a tone, a sound, or a sound wave, or more generally an audio signal, may be received via the audio input 210 and provided as audio input signals 212, 214, and 216, for example, for the audio processing 220, and / or for the reinforcement learning 250, and / or for the neural network 260. For example, the audio signal 212 can be provided for the neural network 260, the audio signal 214 for the reinforcement learning 250, and the audio signal 216 for the audio processing 220 (wherein the audio signals 212, 214, 216 can be the same or can differ in detail (e.g., in the sampling rate, the frequency resolution, the bandwidth, etc.).In this case, the audio signal 212 can be similar to the audio signal 214 and / or the audio signal 216 (or at least describe the same audio content) and can have the same information about the number of frequency channels and frequency bands, so that the audio input signal from the audio input 210 is divided directly, for example without further audio analysis, and can be provided, for example, via multiple outputs or data paths of the audio input 210.

[0050] The audio processing 220 can, for example, have one and / or more parameterized audio processing rules that process / process one or more audio signals 216, for example, such that, based on the incoming audio signal 216 (or the incoming audio signals), a user-customized audio signal 217 is provided (or several user-customized audio signals are provided) using the parameterized audio processing rule, which is parameterized, for example, by the combined parameters 272. The audio processing 220 makes it possible to process the audio input signal 216, which is based on the audio input 210, using the combined parameters 272, for example, using the parameterized audio processing rule, to obtain the user-customized audio signal 217.Optional details and embodiments of the combined parameters 272 will be explained in more detail later in this patent application. Further details and embodiments of the components of the device 200 will follow.

[0051] The audio output 240 can, for example, receive the audio signal 217 modified, reassigned, and customized for the user by the audio processing unit 220 and provide it as a modified or processed audio signal 218 to a coefficient determiner 250 (which is implemented, for example, using reinforcement learning) for determining parameters or coefficients of the processing parameter determination rule (for example, of the neural network 260). Alternatively or additionally, the audio output can, for example, provide the audio signal 217 modified, reassigned, and customized for the user by the audio processing unit 220 as a modified or processed audio signal 219 for an interface, for example, for headphones or loudspeakers, although this is not mandatory.

[0052] Furthermore, embodiments allow additional information of the audio signal 218 to be provided via the audio output 240 to the reinforcement learning 250 (or another device for determining coefficients or parameters of the processing parameter determination rule) in order, for example, to supply a data memory 252 (the content of which may be part of a database) with information about audio signals.

[0053] The audio output signal 218 can, for example, like the audio input signal 214, be provided to the reinforcement learning 250 for determining coefficients or parameters of the processing parameter determination rule 260, so that, for example, the information of the audio input signal 214 and the audio output signal 218 is stored in a data memory 252 as a corresponding database of the device 200.

[0054] In other words, reinforcement learning 250 can, for example, determine coefficients or parameters of processing parameter determination rule 260 using audio signals 218 and 214. Furthermore, reinforcement learning 250 can, for example, expand the database based on audio signals 214, 218 and / or store audio signals 214, 218 in data storage 252. Alternatively or additionally, reinforcement learning can determine at least one user-adapted coefficient 254 or store it in the database.

[0055] It should be noted, however, that the use of the output audio signal 218 by the reinforcement learning 250 (or by another device for determining the coefficients of the processing parameter determination rule, which may replace the reinforcement learning 250) is to be considered optional.

[0056] The database or data memory 252 can include a variety of information, for example, information about the audio input 210 (or about an audio input signal) and / or about one or more of the audio signals 212 and 214 coming from the audio input 210, and / or information about the audio output 240 and / or about the audio signal 218 coming from the audio output 240, and / or information about and for the audio processing 220 and, for example, also at least one user-adapted coefficient 254. User-adapted coefficients 254 can be understood as coefficients that are determined, for example, for use by the processing parameter determination rule 250 based on the database 252 and / or based on a set user parameter 232. However, user-adjusted coefficients can also be understood as audio processing parameters set by the user.

[0057] The coefficients of the processing parameter determination rule, for example edge weights of the neural network, can be based, among other things, on a reinforcement learning method, which is used in the Fig. 2 is identified by the reference numeral 250 as "Reinforcement Learning".

[0058] For example, the reinforcement learning 250 (for example, as a subfunction) can determine the database or the content of the data memory 252 such that the data memory 252 describes an association between different audio input signals 212, 214 and respective user parameters 232 set by the user, for example, a user-adapted coefficient 254.

[0059] For example, by the reinforcement learning 250 determining the database or the content of the data memory 252 such that the data memory 252 (for example additionally) describes an association between the audio output signal 218 and respective user parameters set by the user, for example a user-adapted coefficient 254, coefficients 256 of the neural network can be provided by the reinforcement learning 250 in an advantageous manner.

[0060] Furthermore, the processing parameter determination rule can be designed as a neural network 260 or can be integrated into a neural network to obtain audio processing parameters 262 using, for example, the coefficient 256 determined by reinforcement learning 250. In other words, the neural network 260 can, for example, determine the audio processing parameters 262 based on the audio signal 212 and the coefficient 256 obtained by reinforcement learning 250, so that, as a result, a learning algorithm, for example, sets user-specific audio processing parameters 262.

[0061] The at least one audio processing parameter 262 supplied by the neural network 260 can be a single parameter or can comprise multiple parameters. The neural network 260 can, for example, supply one or more of the following parameters as audio processing parameters 262: a user profile parameter N, and / or a multi-band compression parameter R, and / or a hearing threshold adaptation parameter T, and / or smoothing (or one or more smoothing parameters) and / or compression settings (or one or more compression parameters). Furthermore, one or more parameters can be used (or supplied by the neural network as audio processing parameters 262) for sound adaptation (alternatively or additionally), such as a band-dependent gain G, a noise reduction (or one or more noise reduction parameters) and / or a blind source separation (orone or more parameters of a blind source separation).

[0062] For example, the number of input parameters (for example, of reinforcement learning 250 and / or neural network 260) can be determined as a function of a number C of channels of a multi-channel audio signal, and also as a function of a number B of processing bands, or as a function of a number P of user parameters. For example, the number of user parameters P can be determined as the product of the number of frequency bands B and the number of audio signals or audio channels C.

[0063] Alternatively or additionally, the input parameters (e.g., of reinforcement learning or neural network) may comprise audio features ("Audio Features") N, for example, F=2048 Fourier coefficients per channel for each input (e.g., the audio input signal) and output (e.g., the audio output signal), for example, every 10 ms.

[0064] For example, the number of output parameters (e.g., the output parameters of the neural network 260 or the input parameters of the audio processing) in a learned user profile M can be composed of the number of audio channels (e.g., C), the hearing threshold adaptation T, the multiband compression at rate R, the band-dependent gain G, and two further time constants, wherein the number of values ​​of G, R, T corresponds, for example, to the number of bands B. Furthermore, the value of the learned user profile M (or the values ​​of the learned user profile M) can form the user-adapted coefficient (or parameter) 254 (or a set of user-adapted coefficients or parameters).

[0065] The user control 230 provides at least one user parameter 232, which may include, for example, volume parameters and / or tone control parameters. The user control may, for example, include an interface for visualizing the one or more user parameters.

[0066] A volume control or volume adjustment, which can be performed by the user control 230, can, for example, provide parameters that amplify or attenuate the audio signal. Using a bass control, a treble control, and / or an equalizer, the user can, for example, adjust tone control parameters via the user control 230, which can, for example, be combined as part of the user parameters 232 with the audio processing parameters 262 (provided by the neural network 260) using a combination 270.

[0067] In other words, the user parameters 232 provided by the user control 230 can be combined with the audio processing parameters 262, for example, by addition, multiplication, division, or subtraction. By combining 270 the user parameters 232 with the audio processing parameters 262, for example, combined parameters 272 can be provided to the audio processing 220. Alternatively, the user parameters 232 can also replace the parameters 262, for example, if the user desires a significantly different setting than that specified by the parameters 262.

[0068] In summary, the device 200 processes an audio input signal received via the audio input 210 in the audio processing 220 in order to adapt sound characteristics to the wishes or needs of a user. A processing characteristic of the audio processing 220 is set by the parameters 272, wherein the parameters 272 are influenced, on the one hand, by the neural network 260 and, on the other hand, can be modified by the user via the user control 230. Generally speaking, the reinforcement learning 250 fulfills the function of adapting one or more coefficients (e.g., edge weights) of the neural network such that the parameters provided by the neural network essentially correspond to the user's expectations, i.e., within acceptable tolerances, have the parameter values ​​that the user sets via the user control 230 in respective different acoustic environments.

[0069] This means that after sufficient training in many different acoustic environments, the device can achieve an automatic setting of the audio processing that is comfortable for the user.

[0070] Fig. 3 shows a schematic representation or a schematic block diagram of a device 300 for determining audio processing parameters as a function of an audio input signal and an audio output signal, which is output on the device 200 from the Fig. 2 based.

[0071] It should be noted that in the device 300 according to Fig. 3 Function blocks that are also in the Fig. 2shown, for example, may have similar or identical functionality to corresponding functional blocks in device 200 (but do not necessarily have to). It should further be noted that device 300 may optionally be supplemented with all of the features, functionalities, and details described herein, both individually and in combination.

[0072] The device 300, like the device 200, has an audio input 310 (which can correspond to the audio input 200), an audio processor 320 (which can correspond to the audio processor 220), a user control 330 (which can correspond to the user control 230), an audio output 340 (which can correspond to the audio output 240), a reinforcement learning 350 (which can, for example, correspond to the reinforcement learning 250 in terms of its basic function), a neural network 360 (which can, for example, correspond to the neural network 260 in terms of its basic function) and the combination 370 of the user-specifically set user parameters 332 and the audio processing parameters 362 (which can, for example, correspond to the combination 270).

[0073] Starting from the device 200 from the Fig. 2 The device 300 contains or comprises the Fig. 3additionally an audio analysis 380-1 between the audio input 310 and the neural network 360 and an audio analysis 380-2 between the audio output 340 and the reinforcement learning 350.

[0074] In particular, this arrangement enables the audio analysis 380-1, for example, to receive and analyze the audio input signal 311 emanating from the audio input 310 in order to provide an audio input signal analysis result, for example, information about spectral properties and / or dynamics and / or frequency of the audio input signal 311, in the form of the audio analysis signal 312 and / or 314. The information of the audio analysis result of the audio analysis 380-1 can, for example, be provided to the neural network 360 and the reinforcement learning 350 (for example, simultaneously) via the analyzed audio signals 312, 314.

[0075] The processing parameter determination rule, which may comprise, for example, a part of the neural network 360 (or a part of the reinforcement learning 350) or which is implemented by the neural network 360, may, for example, define a derivation rule for deriving the audio processing parameters 362 from the audio input analysis result. Using the audio analysis 380-1, additional (or compact) information about spectral properties, for example, an intensity value per frequency band and channel, can be obtained in order to provide frequency selectivity, for example, in audio signals (for example, in multi-channel audio signals). Frequency selectivity is required in order to be able to analyze and represent the perceptible sonic aspects of the signal.Generally speaking, the audio analysis 380-1 can significantly reduce the amount of input data to the neural network, for example, compared to a concept in which time-domain samples are input into the neural network. For example, by having the analyzed audio signals 312, 314 contain parameters that describe properties of the audio input signal in a compact form (where a number of parameters per time period is, for example, at least a factor of 10, or at least a factor of 20, or at least a factor of 50 lower than a number of samples per unit of time), the complexity of the neural network 360 can be kept comparatively low. Accordingly, the number of coefficients of the neural network can be kept comparatively low, which facilitates a learning process (for example, through reinforcement learning 350).This is all the more true the better the parameters of the analyzed audio signals are suited to distinguishing different acoustic environments.

[0076] In addition, an audio analysis 380-2 of the audio output signal 342 is performed to provide an audio output signal analysis result for determining the at least one coefficient of the processing parameter rule, for example at least one coefficient of the reinforcement learning 350.

[0077] A "joint" audio analysis of the audio input signal 311 and the audio output signal 342 is also possible (e.g., an audio analysis of both the audio input signal and the audio output signal), whereby separate audio signal analysis results can be provided. Separate in this context means that the audio input signal analysis result can be provided to other components, for example, compared to the audio output signal analysis result. For example, the information from the audio analysis 380-1, 380-2 of the input or output signal can be different from each other or correspondingly identical.

[0078] Embodiments further provide that the audio output 340 provides a modified or processed audio signal 319 for an interface, for example, for headphones or speakers, although this is not mandatory. Furthermore, embodiments allow the audio analysis 380-2 to provide the audio signal 313 for the interface or for another interface. This allows the device 300 to provide the audio signal 319 and 313 to external components, for example, via at least one interface, although this is not mandatory.

[0079] In summary, it can be stated that in device 300, it is not the input audio signal or the output audio signal itself that is fed to the neural network 360 or reinforcement learning 350, but rather one or more corresponding audio analysis results. Thus, by appropriately analyzing the input audio signal and / or the output audio signal in advance, the complexity of the neural network and thus also the complexity of reinforcement learning can be kept low, which significantly reduces the implementation effort.

[0080] Fig. 4 shows a schematic block diagram of a device 400 for determining audio processing parameters in dependence on at least one input signal, which is partly based on the device 200 from the Fig. 2 based.

[0081] It should be noted that in the device 400 according to Fig. 4 Function blocks that are also in the Fig. 2shown, for example, may have similar or identical functionality to corresponding functional blocks in device 200 (but do not necessarily have to). It should further be noted that device 400 may optionally be supplemented with all of the features, functionalities, and details described herein, both individually and in combination.

[0082] The device 400 comprises an audio input 410 (which may, for example, correspond to audio input 210), an audio processor 420 (which may, for example, correspond to audio processor 220), a user controller 430 (which may, for example, correspond to user controller 230), an audio output 440 (which may, for example, correspond to audio output 240), a reinforcement learning 450 (which may, for example, correspond to reinforcement learning 250 in terms of its basic function), a neural network 460 (which may, for example, correspond to neural network 260 in terms of its basic function), a combination 470 (which may, for example, correspond to combination 270), and an audio analysis 480 (which may, for example, correspond to audio analysis 380-1) between the audio input 410 and the neural network 460 and the reinforcement learning 450.

[0083] Compared to device 300, device 400 does not include audio analysis of audio output 440, and compared to device 200, no audio output signal from audio output 440 is provided to reinforcement learning 450. In other words, reinforcement learning 450 does not receive any information about the audio output signal.

[0084] Instead, reinforcement learning 450 is based on the combined parameters 472, 473, or on information 433 describing changes or adjustments by the user to the audio processing parameters 462 provided by the neural network 460. Furthermore, reinforcement learning uses the audio input signal analysis result 414.

[0085] In other words, reinforcement learning 450 can determine a database 452 depending on the user parameters set by the user or the combined parameters 472, 473, such that entries in database 452 represent the user parameters 472, 473 set by the user. Database 452 can be provided or used to determine the coefficients 456 of the processing parameter determination rule, or of the neural network 460. This allows a prediction model to be determined that is directly based on user parameters (or the audio signal processing parameters 472 adjusted by the user), which are directly assigned to reinforcement learning 450.

[0086] Optionally, the one or more combined parameters 472, 473 or user parameters can also be fed directly into the neural network 460 during operation by means of the combined parameter 474, so that, for example, the compressor settings and / or other parameters for the audio processing parameters 462 can be provided as output.

[0087] Alternatively or optionally, the respective user parameters 432 set by the user can be provided directly to reinforcement learning 450 (as shown at reference numeral 433), although this is not mandatory. For example, information about how much the user changes the parameters 462 provided by the neural network 460 can be used for reinforcement learning. If the user changes the parameters 462 provided by the neural network 460 only slightly or not at all, it can be assumed that the user is completely or at least highly satisfied with the current functionality of the neural network, so that the coefficients of the neural network need not be changed at all or only slightly.If, however, the user makes significant changes to the parameters 462, reinforcement learning can assume that a significant change in the coefficients of the neural network is necessary to ensure that the parameters 462 provided by the neural network correspond to the user's expectations. Therefore, for example, information 433 describing a user intervention can be used by reinforcement learning to trigger learning and / or determine the extent of the changes to the coefficients of the neural network.

[0088] Overall, the device allows according to the Fig. 4 to learn and / or (e.g., continuously) improve the coefficients 456 of the neural network 460 in an efficient manner.

[0089] Fig. 5shows a device 500 which has similar properties to the devices 200, 300 and 400. It should be noted that in the device 500 according to Fig. 5 Function blocks that are also included in the Fig. 2 , Fig. 3 and Fig. 4 shown, for example, may have (but do not necessarily have to have) similar or identical functionality to corresponding functional blocks in device 200, device 300, and device 400. It should further be noted that device 500 may optionally be supplemented with all of the features, functionalities, and details described herein, both individually and in combination.

[0090] The schematic block diagram of the Fig. 5shows the device 500, comprising an audio input 510 (which can correspond, for example, to the audio input 210), an audio processor 520 (which can correspond, for example, to the audio processor 220), a user control 530 (which can correspond, for example, to the user control 230), an audio output 540 (which can correspond, for example, to the audio output 240), a reinforcement learning 550 (which, for example, can correspond in terms of its basic function to the reinforcement learning 250), a neural network 560 (which, for example, can correspond in terms of its basic function to the neural network 260) and a combination 570 (which, for example, can correspond to the combination 270).

[0091] For example, the device 500 does not include any audio analysis of the audio input signal or any audio analysis of the audio output signal, so that the audio signals 512 and 514 can be passed directly from the audio input 510 to the reinforcement learning 550 and the neural network 560, respectively. Optionally, however, the device 500 can also perform an audio analysis of the audio input signal.

[0092] As already mentioned in the Fig. 2 As mentioned above, an audio input signal 512 may be provided to the neural network 560 and an audio input signal 514 may be provided to the reinforcement learning 550. In contrast to the device 400, the reinforcement learning 550 of the device 500 may be based on the audio input signal 514 and the one or more audio processing parameters 572 provided to the audio processing 520 (or actually used by the audio processing 520).

[0093] Optionally, the user parameter, or the combined parameter 572, can be provided to the neural network 560, so that the user parameter 572 and the coefficient(s) provided by the reinforcement learning 550 are received or provided as input variables of the neural network 560.

[0094] The device 500 allows a particularly efficient adjustment of the coefficients of the neural network, since the reinforcement learning 550 takes into account the parameters actually used by the audio signal processing 520 and can thus determine or optimize the coefficients of the neural network very precisely.

[0095] Fig. 6shows a schematic flow diagram of a method 600 for operating a device, such as device 100, 200, 300, 400, or 500, or more generally for obtaining audio processing parameters. A first step 610 comprises a user-specific determination of at least one coefficient of a processing parameter determination rule based on audio signals obtained during user operation. A second step 620 comprises obtaining audio processing parameters using the processing parameter determination rule based on the audio input signal.

[0096] Method 600 is implemented, for example, such that audio processing parameters are determined as a function of at least one audio input signal. Method 600 can be implemented such that sound processing or audio processing based on directly recorded ambient noise (where, for example, an audio input signal leads to an adjustment of audio processing parameters) leads to an improvement in the individual perception of sound. For example, it can be achieved that the coefficients of the processing parameter determination rule are based on audio input signals obtained during user operation and are determined user-individually (for example, in real time), so that audio processing parameters are obtained using a neural network whose coefficients are determined by reinforcement learning or even continuously adjusted based on the audio input signal.

[0097] The method 600 may optionally be supplemented with any features, functionalities, and details described herein, even if described with respect to devices. The method may be supplemented with these features, functionalities, and details both individually and in combination. Further examples

[0098] In the following, some aspects of the present invention are described which can be used individually or in combination in embodiments.

[0099] Situation-dependent control parameters that can be set by the user, or user parameters that can be set by the user, can be determined, for example, by analyzing the incoming and outgoing audio signal, as in the Fig. 3 shown, the sound processing can be integrated into the training of a self-reinforcing learning algorithm.

[0100] The incoming audio signal can contain the sound environment. This allows changes in the sound environment and user settings to be learned inherently, for example, in runtime.

[0101] The self-reinforcing learning algorithm can, for example, develop a user-specific set of parameters for sound processing from this data. This automatically applies the same control parameters under the same external conditions, but also allows for further user adjustments in the situation itself, which are integrated into the learning system (e.g., based on a reinforcement learning principle). For example, the machine learning system and the application can adapt to the user's sound preferences in a continuous learning process. Algorithms such as those used in hearing aids can be integrated and controlled for sound adaptation. These can include, for example, multiband compression with rate R and hearing threshold adjustment T and band-dependent gain G, noise reduction, or blind source separation.

[0102] The incoming audio signal, the sound processing parameters, and / or the audio signal processed with the sound processing parameters can be stored in a cloud (e.g., a central data storage), for example, to train the user profile. At the same time, the sound processing parameters selected by the user, or user parameters, can be applied to the incoming audio signal. The number of input parameters for reinforcement learning, e.g., a CNN (convolutional neural network), can, for example, consist of multi-channel audio input (e.g., with C=4 channels) and audio output (e.g., with C=2 channels). The number of output parameters in the learned parameter set M can, for example, consist of M = C * (T + R + G) + 2 time constants, where the number of values ​​of G, R, T can, for example, correspond to the number of processing bands B (e.g., B = 8).

[0103] In the following, some aspects of the present invention are described which can be used individually or in combination in embodiments.

[0104] One possible implementation of the method or device in the field of sound control is, for example, that a user wears a sound reproduction device (e.g. a hearable or an earphone with additional function) that is equipped with a system with integrated sound amplification and audio analysis, for example as in the Fig. 3 or the Fig. 4shown. The user can control the sound amplification parameters, for example, with an app (or application software), for example using the user control described above. In the background, the audio analysis can, for example, continuously record and analyze the auditory scene using microphony and evaluate it, for example, in terms of dynamics and / or frequency and / or spectral properties (for example in the audio analysis). In a specific auditory scene, e.g. while driving in a car on the highway, the user can perform a sound adjustment using an app and thus change the sound amplification parameters (for example, parameters 272).The system (for example, reinforcement learning 250) can establish an algorithmic relationship between the user's parameter changes and the analysis of the auditory scene. From this, it can develop a predictive model that integrates further ad hoc sound adjustments by the user using artificial intelligence (AI) (and describes them, for example, using coefficients 256). This means that individualized AI control (AI here means, for example, artificial intelligence) or individualized AI control (artificial intelligence, AI) is enabled or provided by the device.

[0105] If, for example, the user returns to the same auditory scene at a different time—in this case, a moving car on the highway—the predictive model is applied, and the sound amplification parameters (e.g., parameters 262) are automatically set or provided by the system (e.g., by the neural network 260 defined by coefficients 256). If the user makes further sound adjustments (e.g., via interface 230), these can be integrated ad hoc into the self-learning system.

[0106] In the following, some aspects of the present invention are described, which can be applied individually or in combination in embodiments, and which, for example, represent differences to the Github publication "liketohear-ai-pt". According to one (optional) aspect of the invention, the prediction model is based on a multidimensional optimization in real time that analyzes all available frequency bands. According to one (optional) aspect of the invention, for example, reinforcement learning methods and unsupervised learning methods are used. According to one (optional) aspect of the invention, the adaptation (or adaptations), for example, of the processing parameter determination rule and / or the audio processing parameters, can take place continuously at runtime.

[0107] In the following, some aspects of the present invention are described which can be used individually or in combination in embodiments which, for example, represent differences from the published patent application US 2015 195641 A1.

[0108] Embodiments according to the invention, for example, primarily relate to intuitive and ergonomic user control of sounds in everyday acoustic environments and therefore prefer generalized setting options for the following reasons: Dividing signals into individual "types of sounds" in real time is hardly feasible in everyday acoustic situations. Therefore, the present invention does not apply this method, but rather covers a multitude of sound possibilities with a 2-dimensional parameter space. With signal separation, user settings would have to be made separately for each object and each contextual situation. In everyday acoustic environments with rapidly changing listening situations, user control becomes too complex and therefore not ergonomically applicable. With the present invention, the user can perform complex sound adjustments (e.g., in device 230) using a simple and intuitive interface, such as a 2D touch surface of a smartphone. The acoustic properties of individual sounds could sound different when combined than in preference. e.g.Sounds like music as a foreground or background noise. Therefore, with the present invention, for example, the complexity of the auditory scene is adapted to an optimized perception of all existing sounds for the user. Settings for individual signals do not adapt dynamically to changing environmental conditions. For example, with softly spoken speech or softly played music, even a slight increase in the volume of the background noise can make speech unintelligible or the music inaudible.

[0109] In the following, some aspects of the present invention are described which can be used individually or in combination in embodiments which, for example, represent differences from the published patent application US 2020 0066264 A1.

[0110] In the published patent application US 2020 0066264 A1, a processor controls the sound processing of the hearing aid based on "user preferences and interests" and "historical activity patterns".

[0111] In embodiments of the present invention, however, the sound processing of the hearing aid is carried out, for example, on the basis of external conditions measured in real time, for example as in the Fig. 2 represented, controlled.

[0112] In summary, according to one aspect of the invention, the above-mentioned criteria or requirements are integrated into a learning method or device that learns from user settings in real time and applies them automatically to improve the user's individual perception of sound or tones in the form of audio signals. By means of the present invention, signal reproduction or audio reproduction optimized to user preferences can be realized.

[0113] Thus, according to one aspect of the present invention, it can be taken into account that the individual perception of sound and thus the individual requirements for the sound or euphony in their adaptation of sound reproduction devices differ, among other things, according to the following criteria: Individuality Situational needs External conditions

[0114] According to one aspect of the invention, embodiments according to the invention may take into account that sound perception differs from person to person.

[0115] For example, a conversation with someone in a room full of people and with a loud background noise may be more difficult for one person than for another. Likewise, the same sound reproduction setting may be perceived differently depending on individual needs.

[0116] According to one aspect of the invention, embodiments according to the invention may take into account that environmental parameters, such as the auditory environment, also significantly influence the control values ​​for sound adaptation of a sound reproduction device.

[0117] In summary, embodiments according to the present invention provide a device and a method that perform sound processing based on ambient noise that is directly recorded or measured. Based on these recordings and the user parameters set by the user, a learning algorithm, for example, generates a predictive model that allows for further adjustments in the situation itself, which are integrated into the learning system to improve the user's individual perception of sound or tones in the form of audio signals.

[0118] Although some aspects have been described in connection with a device, it is understood that these aspects also represent a description of the corresponding method, so that a block or component of a device can also be understood as a corresponding method step or as a feature of a method step. Similarly, aspects described in connection with or as a method step also represent a description of a corresponding block, detail, or feature of a corresponding device.

[0119] Depending on specific implementation requirements, embodiments of the invention may be implemented in hardware or software. The implementation may be performed using a digital storage medium, such as a floppy disk, a DVD, a Blu-ray Disc, a CD, a ROM, a PROM, an EPROM, an EEPROM, or a FLASH memory, a hard disk, or other magnetic or optical storage device storing electronically readable control signals that can interact or cooperate with a programmable computer system to perform the respective method. Therefore, the digital storage medium may be computer-readable.Some embodiments according to the invention thus comprise a data carrier having electronically readable control signals capable of interacting with a programmable computer system such that one of the methods described herein is carried out.

[0120] In general, embodiments of the present invention can be implemented as a computer program product with a program code, wherein the program code is effective to perform one of the methods when the computer program product is run on a computer. The program code can also be stored, for example, on a machine-readable medium.

[0121] Other embodiments include the computer program for performing one of the methods described herein, wherein the computer program is stored on a machine-readable carrier.

[0122] In other words, one embodiment of the method according to the invention is thus a computer program comprising program code for performing one of the methods described herein when the computer program is run on a computer. Another embodiment of the method according to the invention is thus a data carrier (or a digital storage medium or a computer-readable medium) on which the computer program for performing one of the methods described herein is recorded.

[0123] A further embodiment of the method according to the invention is thus a data stream or a sequence of signals that represents the computer program for carrying out one of the methods described herein. The data stream or the sequence of signals can be configured, for example, to be transferred via a data communication connection, for example, via the Internet.

[0124] A further embodiment comprises a processing device, for example a computer or a programmable logic device, which is configured or adapted to carry out one of the methods described herein.

[0125] A further embodiment comprises a computer on which the computer program for performing one of the methods described herein is installed.

[0126] In some embodiments, a programmable logic device (e.g., a field-programmable gate array, an FPGA) may be used to perform some or all of the functionalities of the methods described herein. In some embodiments, a field-programmable gate array may interact with a microprocessor to perform any of the methods described herein. In general, in some embodiments, the methods are performed by any hardware device. This may be general-purpose hardware such as a computer processor (CPU) or method-specific hardware such as an ASIC.

[0127] The above-described embodiments are merely illustrative of the principles of the present invention. It is understood that modifications and variations of the arrangements and details described herein will be apparent to others skilled in the art. Therefore, it is intended that the invention be limited only by the scope of the following claims and not by the specific details presented in the description and explanation of the embodiments herein.

Claims

1. An apparatus (100; 200; 300; 400; 500) for determining audio processing parameters (120; 262; 362; 462; 562) in dependence on at least one audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516); wherein the apparatus (100; 200; 300; 400; 500) is configured to determine at least one coefficient (142; 256; 356; 456; 556) of a processing parameter determination rule (140; 250; 350; 450; 550) in a user-individual manner based on audio signals (217, 218, 219; 313, 317, 318, 319, 342; 417; 517) obtained during user operation; wherein the apparatus (100; 200; 300; 400; 500) is configured to obtain the audio processing parameters (120; 262; 362; 462; 562) by using the processing parameter determination rule (140; 250; 350; 450; 550) based on the audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516); wherein the apparatus is configured to determine a database (252; 352; 452; 552) in dependence on the at least one audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516) such that entries of the database (252; 352; 452; 552) describe the audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516); wherein the apparatus is configured to determine the database (252; 352; 452; 552) in dependence on an audio output signal (218, 219, 313, 318, 319, 342), which is obtained in dependence on a user parameter such that entries of the database (252; 352; 452; 552) describe the audio output signal (218, 219, 313, 318, 319, 342); wherein the apparatus is configured to adapt the at least one coefficient (142; 256; 356; 456; 556) of the processing parameter determination rule (140; 250; 350; 450; 550) based on the database (252; 352; 452; 552) acquired by the apparatus in order to adapt the processing parameter determination rule (140; 250; 350; 450; 550) in a user-individual manner in order to obtain audio processing parameters (120; 262; 362; 462; 562) that are adapted in a user-individual manner.

2. The apparatus (100; 200; 300; 400; 500) according to claim 1, wherein the apparatus is configured to determine the database (252; 352; 452; 552) in dependence on user parameters (232; 332; 432, 433; 532) adjusted by the user such that entries of the database (252; 352; 452; 552) describe the user parameters (232; 332; 432, 433; 532) adjusted by the user.

3. The apparatus (100; 200; 300; 400; 500) according to claim 1 or 2, wherein the apparatus is configured to determine the database (252; 352; 452; 552) such that the database (252; 352; 452; 552) describes an allocation between different audio input signals (110,112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516) and respective user parameters (232; 332; 432, 433; 532) adjusted by the user; and / or wherein the apparatus is configured to determine the database (252; 352; 452; 552) such that the database (252; 352; 452; 552) describes an allocation between different audio output signals (218, 219, 313, 318, 319, 342) and respective user parameters (232; 332; 432, 433; 532) adjusted by the user.

4. The apparatus (100; 200; 300; 400; 500) according to any of the preceding claims, wherein the apparatus is configured to determine and / or adapt the at least one coefficient (142; 256; 356; 456; 556) of the processing parameter determination rule (140; 250; 350; 450; 550) based on at least one audio processing parameter (120; 262; 362; 462; 562) corrected and / or amended by a user.

5. The apparatus (100; 200; 300; 400; 500) according to any of the preceding claims, wherein the apparatus is configured to perform audio processing (220; 320; 420; 520) based on the audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514,516) and based on the audio processing parameter (120; 262; 362; 462; 562) in order to obtain the audio signals (217, 218, 219; 313, 317, 318, 319, 342) that are adapted in a user-individual manner.

6. The apparatus (100; 200; 300; 400; 500) according to any of the preceding claims, wherein the apparatus is configured to determine the coefficients (142; 256; 356; 456; 556) of the processing parameter determination rule (140; 250; 350; 450; 550) by using a comparison of the audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516) and an audio output signal (218, 219, 313, 318, 319, 342) provided by the audio processing (220; 320; 420; 520) by using the audio processing parameters (120; 262; 362; 462; 562).

7. The apparatus (100; 200; 300; 400; 500) according to any of the preceding claims, wherein the apparatus is configured to provide the user parameters (232; 332; 432, 433; 532) adjusted by the user as output quantity instead of the audio processing parameters (120; 262; 362; 462; 562), and wherein the user parameters (232; 332; 432, 433; 532) adjusted by the user include volume parameters and / or sound parameters and / or equalizer parameters; or wherein the apparatus is configured to combine the user parameters (232; 332; 432, 433; 532) with the audio processing parameters (120; 262; 362; 462; 562) to obtain combined parameters (272; 372; 472, 473, 474; 572, 573) of the audio processing (220; 320; 420; 520) and provide the same as output quantity.

8. The apparatus (100; 200; 300; 400; 500) according to any of the preceding claims, wherein the apparatus is configured to perform audio analysis of the audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516) to provide an audio input signal analysis result for determining the at least one coefficient (142; 256; 356; 456; 556) of a processing parameter determination rule (140; 250; 350; 450; 550); and / or wherein the apparatus is configured to perform audio analysis of the audio output signal (342) to provide an audio output signal analysis result for determining the at least one coefficient (142; 256; 356; 456; 556) of the processing parameter determination rule (140; 250; 350; 450; 550).

9. The apparatus (100; 200; 300; 400; 500) according to any of the preceding claims, wherein the audio processing parameters (120; 262; 362; 462; 562) include at least one multiband compression parameter R and / or at least one hearing threshold adaptation parameter T and / or at least one band-dependent amplification parameter G and / or at least one disturbing noise reduction parameter and / or at least one blind source separation parameter and / or at least one sound direction parameter and / or at least one binaural parameter and / or at least one parameter of adaptive filters.

10. The apparatus (100; 200; 300; 400; 500) according to any of the preceding claims, wherein the apparatus includes a neuronal network (260; 360; 460; 560) configured to obtain the audio processing parameters (120; 262; 362; 462; 562) by using the processing parameter determination rule (140; 250; 350; 450; 550).

11. The apparatus (100; 200; 300; 400; 500) according to any of the preceding claims, wherein the apparatus is configured to provide and / or adapt the processing parameter determination rule (140; 250; 350; 450; 550) based on a method of consolidating learning and / or based on a method of reinforcement learning and / or based on a method of unsupervised learning and / or based on a method of multivariate prediction and / or based on a multidimensional parameter space determined with multivariable regression in order to determine the audio processing parameter (120; 262; 362; 462; 562).

12. The apparatus (100; 200; 300; 400; 500) according to any of the preceding claims, wherein the apparatus is configured to determine the at least one coefficient (142; 256; 356; 456; 556) of the processing parameter determination rule (140; 250; 350; 450; 550) in a user-individual manner during user operation in order to obtain the audio processing parameters (120; 262; 362; 462; 562) in real time and / or to determine and / or adapt the amended audio processing parameters (120; 262; 362; 462; 562) in real time.

13. A hearing aid, wherein the hearing aid comprises audio processing; and wherein the hearing aid comprises an apparatus for determining audio processing parameters according to any of claims 1 to 12, wherein the audio processing is configured to process an audio input signal in dependence on the audio processing parameters.

14. A method (600) for determining audio processing parameters in dependence on at least one audio input signal, the method comprising: determining, in a user-individual manner, at least one coefficient of a processing parameter determination rule based on audio signals obtained during user operation; and obtaining audio processing parameters by using the processing parameter determination rule based on the audio input signal; wherein a database (252; 352; 452; 552) is determined in dependence on the at least one audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516) such that entries of the database (252; 352; 452; 552) describe the audio input signal (110, 112; 212, 214, 216; 311, 316; 411, 416; 512, 514, 516); wherein the database (252; 352; 452; 552) is determined in dependence on an audio output signal (218, 219, 313, 318, 319, 342), which is obtained in dependence on a user parameter such that entries of the database (252; 352; 452; 552) describe the audio output signal (218, 219, 313, 318, 319, 342); wherein the at least one coefficient (142; 256; 356; 456; 556) of the processing parameter determination rule (140; 250; 350; 450; 550) is adapted based on the database (252; 352; 452; 552) acquired by the apparatus in order to adapt the processing parameter determination rule (140; 250; 350; 450; 550) in a user-individual manner in order to obtain audio processing parameters (120; 262; 362; 462; 562) that are adapted in a user-individual manner.

15. A computer program comprising program code for performing the method according to claim 14 when the program runs on a computer.