MUSIC CLASSIFIERS AND RELATED METHODS

A computationally efficient music classifier for hearing aids transforms audio signals into frequency bands, using parallel decision-making units and neural networks to detect music, ensuring accurate adjustments without increasing power consumption or size.

DE102019004239B4Active Publication Date: 2026-01-22SEMICON COMPONENTS IND LLC
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
DE102019004239
Authority / Receiving Office
DE · DE
Patent Type
Patents
Current Assignee / Owner
Priority Date
2019-06-03
Filing Date
2019-06-14
Publication Date
2026-01-22
Estimated Expiration
2039-06-14

AI Technical Summary

Technical Problem

Hearing aids face challenges in efficiently detecting music without significantly increasing power consumption or size, as automated audio adjustments for enhanced user experience require complex computations.

Method used

A computationally efficient music classifier for hearing aids that transforms audio signals into frequency bands, uses parallel decision-making units to evaluate features, and combines ratings to determine music presence, incorporating modulation activity tracking and neural networks for accurate detection.

Benefits of technology

The music classifier provides accurate music detection with low power consumption, allowing hearing aids to adjust audio processing in real-time without affecting battery life or size.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

Music classifier for an audio device, the music classifier comprising the following: a signal processing unit configured to transform a digitized time-domain audio signal into a corresponding frequency-domain signal encompassing a variety of frequency bands; a multitude of decision-making units operating in parallel, each configured to evaluate one or more of the multitude of frequency bands in order to determine a multitude of feature ratings, each feature rating corresponding to a property associated with music; and a combination and music recording unit configured to to receive the multitude of feature ratings asynchronously from the decision-making units, and to combine the multitude of feature ratings over a period of time to determine whether the audio signal includes music, wherein the multitude of decision-making units includes a modulation activity tracking unit configured to detect broadband modulation based on a minimum averaged energy and a maximum averaged energy of a sum of the multitude of frequency bands.
Need to check novelty before this filing date? Find Prior Art

Description

CROSS-REFERENCE TO RELATED REGISTRATIONS

[0001] This application claims the benefits of the preliminary US application No. 62 / 688 726, filed on June 22, 2018, entitled “A COMPUTATIONALLY EFFICIENT SUB-BAND MUSIC CLASSIFIER”, which is hereby incorporated in its entirety by reference.

[0002] This application relates to the non-preliminary application No. 16 / 375 039, filed on April 4, 2019, entitled “COMPUTATIONALLY EFFICIENT SPEECH CLASSIFIER AND RELATED METHODS”, which claims priority over the preliminary US application No. 62 / 659 937, filed on April 19, 2018, both of which are incorporated herein by reference in their entirety. AREA OF REVELATION

[0003] This disclosure relates to a device for music detection and associated methods for music detection. In particular, this disclosure relates to detecting the presence or absence of music in applications with limited processing power, such as hearing aids. BACKGROUND

[0004] Hearing aids can be adjusted to process audio differently based on the environment type and / or the type of audio a user wishes to experience. It may be desirable to automate this adjustment to provide a more natural experience. Automation can involve detecting (i.e., classifying) the environment type and / or audio type. However, this detection can be computationally complex, implying that a hearing aid with automated adjustment will consume more power than one with manual (or no) adjustment. Power consumption can increase further as the number of detectable environment and / or audio types is increased to enhance the user's natural experience.Since, in addition to providing a natural experience, it is highly desirable for a hearing aid to be small and to operate for extended periods on a single charge, there is a need for an environment-type and / or audio-type sensor that operates accurately and efficiently without significantly increasing the power consumption and / or size of the hearing aid.

[0005] US 2017 / 0180875A1 describes a method for operating a hearing aid system based on a classification of the current sound environment. SUMMARY

[0006] The problem is solved by a music classifier for an audio device having the features of claim 1, a method for detecting music in an audio signal of claim 6, and a hearing aid of claim 9. The music classifier includes a signal conditioning unit configured to transform a digitized time-domain audio signal into a corresponding frequency-domain signal encompassing a plurality of frequency bands. The music classifier also includes a plurality of decision-making units operating in parallel, each configured to evaluate one or more of the plurality of frequency bands to determine a plurality of feature ratings, each feature rating corresponding to a property (i.e., a feature) associated with music.The music classifier also includes a combination and music detection unit configured to combine feature ratings over a period of time to determine whether the audio signal contains music.

[0007] In possible implementations, the decision-making units of the music classifier comprise a modulation activity tracking unit and may include one or more of a bar detection unit and a pitch detection unit.

[0008] In one possible implementation, the clock detection unit can detect a repeating clock pattern in a first (e.g., lowest) frequency band of the multitude of frequency bands based on a correlation, while in another possible implementation, the clock detection unit can detect the repeating pattern based on an output of a neural network that receives the multitude of frequency bands as its input.

[0009] In one possible implementation, the combination and music detection unit is configured to apply a weight to each feature rating to obtain weighted feature ratings and to sum these weighted feature ratings to obtain a music rating. This possible implementation can further be characterized by accumulating music ratings for a multitude of frames and calculating an average of these music ratings across the multitude of frames. This average music rating can then be compared to a threshold to determine whether or not music is present in the audio signal. In another possible implementation, hysteresis control can be applied to the output of this threshold comparison, making the music / no-music decision less susceptible to erroneous changes (e.g., due to noise).In other words, the final determination of the current state of the audio signal (i.e., music / no music) can be based on a previous state (i.e., music / no music) of the audio signal. In another possible implementation, the combination and music detection approach described above is replaced by a neural network that receives the feature ratings as inputs and provides an output signal indicating either a music state or a state without music.

[0010] In another aspect, the present disclosure generally describes a method for music capture. In the method, an audio signal is received and digitized to obtain a digitized audio signal. The digitized audio signal is converted into a plurality of frequency bands. The plurality of frequency bands is then applied to a plurality of decision-making units operating in parallel to generate corresponding feature evaluations. Each feature evaluation corresponds to a probability that a particular musical feature (e.g., a beat, a tone, high modulation activity, etc.) is contained in the audio signal (i.e., based on data from one or more frequency bands). Finally, the method includes combining the feature evaluations to capture music in the audio signal.

[0011] In one possible implementation, an audio device (e.g., a hearing aid) performs the procedure described above. For example, a non-volatile, computer-readable medium containing computer-readable instructions can be executed by a processor of the audio device to cause the audio device to perform the procedure described above.

[0012] In another aspect, the present disclosure generally describes a hearing aid. The hearing aid includes a signal processing stage configured to convert a digitized audio signal into a variety of frequency bands. The hearing aid further includes a music classifier coupled to the signal processing stage. The music classifier includes a feature detection and tracking unit, which in turn includes a variety of decision-making units operating in parallel. Each decision-making unit is configured to generate a feature score corresponding to the probability that a particular musical feature is included in the audio signal. The music classifier also includes a combination and music detection unit configured, based on the feature score from each decision-making unit, to detect music in the audio signal.The combination and music capture unit is further configured to generate a first signal indicating music while music is being captured in the audio signal, and is configured to generate a second signal indicating no other music signal.

[0013] In one possible implementation, the hearing aid includes an audio signal modifier stage coupled with the signal processing stage and the music classifier. The audio signal modifier stage is configured to process the multiple frequency bands differently when a music signal is received compared to when no music signal is received.

[0014] The foregoing illustrative summary, as well as other exemplary aims and / or benefits of the disclosure and the manner in which they are achieved, are further explained in the following detailed description and in the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS Fig. Figure 1 is a functional block diagram, which generally represents an audio device with a music classifier according to one possible implementation of the present disclosure. Fig. Figure 2 is a block diagram that generally represents a signal processing stage of the audio device. Fig. 1 represents. Fig. Figure 3 is a block diagram that generally represents a feature capture and tracking unit of the music classifier. Fig. 1 shows. Fig. 4A is a block diagram that generally represents a beat acquisition unit of the feature acquisition and tracking unit of the music classifier according to a first possible implementation. Fig. 4B is a block diagram that generally represents a beat acquisition unit of the feature acquisition and tracking unit of the music classifier according to a second possible implementation. Fig. Figure 5 is a block diagram that generally represents a tone acquisition unit of the feature acquisition and tracking unit of the music classifier according to one possible implementation. Fig. Figure 6 is a block diagram that generally represents a modulation and activity tracking unit of the feature acquisition and tracking unit of the music classifier according to one possible implementation. Fig. Figure 7A is a block diagram that generally represents a combination and music acquisition unit of the music classifier according to a first possible implementation. Fig. 7B is a block diagram that generally represents a combination and music capture unit of the music classifier according to a second possible implementation. Fig. Figure 8 is a hardware block diagram, which generally shows an audio device according to one possible implementation of the present disclosure. Fig. 9 is a method for capturing music in an audio device according to a possible implementation of the present disclosure.

[0015] The components in the drawings are not necessarily to scale with respect to each other. The same reference symbols denote corresponding parts in the different views. DETAILED DESCRIPTION

[0016] The present disclosure relates to an audio device (i.e., a device) and an associated method for music classification (e.g., music detection). As discussed herein, music classification (music detection) refers to identifying musical content within an audio signal, which may include other audio content such as speech and noise (e.g., background noise). Music classification may include identifying music within an audio signal so that the audio signal can be appropriately modified. For example, the audio device may be a hearing aid that may include algorithms for noise reduction, feedback cancellation, and / or audio bandwidth control. These algorithms can be activated, deactivated, and / or modified based on the detection of music.For example, a noise reduction algorithm can decrease signal attenuation levels while capturing music to preserve its quality. In another example, a feedback suppression algorithm can be prevented (essentially blocked) from suppressing tones of music, as doing so would otherwise suppress a feedback signal. In yet another example, the bandwidth of audio signals presented to a user by the audio device, which is normally low to conserve power, can be increased when music is present to enhance the listening experience.

[0017] The implementations described herein can be used to implement a computationally efficient and / or performance-efficient music classifier (and associated procedures). This can be achieved by using decision-making units, each capable of capturing a property (i.e., features) corresponding to music. Each decision-making unit alone may not be able to classify music with high accuracy. However, the outputs of all decision-making units can be combined to form an accurate and robust music classifier. One advantage of this approach is that the complexity of each decision-making unit can be limited to save performance without negatively impacting the overall performance of the music classifier.

[0018] The exemplary implementations described herein detail various operating parameters and techniques, such as thresholds, weights (coefficients), calculations, rates, frequency ranges, frequency bandwidths, and so on. These exemplary operating parameters and techniques are given as examples, and the specific operating parameters, values, and techniques (e.g., computational approaches) used depend on the particular implementation. Furthermore, various approaches to determining the specific operating parameters and techniques for a given implementation can be used in a number of ways, such as using empirical measurements and data, using training data, and so forth.

[0019] Fig. Figure 1 is a functional block diagram, generally representing an audio device implementing a music classifier. As shown in Fig. As shown in Figure 1, the audio device 100 includes an audio transducer (e.g., a microphone 110). The analog output of the microphone 110 is digitized by an analog-to-digital (A / D) converter 120. The digitized audio signal is modified for processing by a signal conditioning stage 130. For example, the time-domain audio signal represented by the digitized output of the A / D converter 120 can be converted by the signal conditioning stage 130 into a frequency-domain representation, which can then be modified by an audio signal modifier stage 150.

[0020] The audio signal modifier stage 150 can be configured to improve the quality of the digital audio signal by removing noise, filtering, amplifying, and so forth. The processed audio signal (e.g., improved quality) can then be transformed into a digital time-domain signal 151 and converted into an analog signal by a digital-to-analog (D / A) converter 160 for playback on an audio output device (e.g., the loudspeaker 170) to generate output audio signals 171 for a user.

[0021] In some possible implementations, the audio device 100 is a hearing aid. The hearing aid receives audio signals (i.e., sound pressure waves) from an environment 111, processes the audio signals as described above, and presents (e.g., using a receiver (i.e., a loudspeaker) of a hearing aid 170) the processed version of the audio signals as output audio signals 171 (i.e., sound pressure waves) to a user wearing the hearing aid. An audio signal modification stage implemented by algorithms can help a user understand speech and / or other sounds in the user's environment. Furthermore, it may be advantageous for the selection and / or setting of these algorithms to be automatic based on different environments and / or sounds. Accordingly, the hearing aid can implement one or more classifiers to detect different environments and / or sounds.The output of one or more classifiers can be used to automatically set one or more functions of the audio signal modifier stage 150.

[0022] One aspect of desirable operation might be characterized by the fact that one or more classifiers provide highly accurate results in real time (as perceived by a user). Another aspect of desirable operation might be characterized by low power consumption. For example, a hearing aid and its normal operation might define a certain size and / or time between charging of an energy storage unit (e.g., accumulator or battery). Accordingly, it is desirable that automatic modification of the audio signal based on the real-time operation of one or more classifiers does not significantly affect the size and / or time between battery changes for the hearing aid.

[0023] The in Fig. The audio device 100 shown includes a music classifier 140 configured to receive signals from the signal conditioning stage 130 and produce an output corresponding to the presence and / or absence of music. For example, when music is detected in audio signals received by the audio device 100, the music classifier 140 can output a first signal (e.g., a logic high signal). When no music is detected in audio signals received by the audio device, the music classifier can output a second signal (e.g., a logic low signal). The audio device can further include one or more other classifiers 180 that output signals based on other conditions. For example, the classifier described in US 2019 / 0325899A1 can be included in one or more of the other classifiers 180 in a possible implementation.

[0024] The music classifier 140 disclosed herein receives as its input the output of a signal processing stage 130. The signal processing stage can also be used as part of the routine audio processing for the hearing aid. Accordingly, one advantage of the disclosed music classifier 140 is that it can use the same processing as other stages, thus saving complexity and reducing performance requirements. Another advantage of the disclosed music classifier is its modularity. The audio device can deactivate the music classifier without affecting its normal operation. In one possible implementation, for example, the audio device could deactivate the music classifier 140 upon detecting a low-energy state (i.e., a low battery).

[0025] The audio device 100 includes stages (e.g., signal conditioning 130, music classifier 140, audio signal modification 150, signal transformation 151, other classifiers 180) that can be implemented as hardware or as software. For example, the stages can be implemented as software running on a general-purpose processor (e.g., CPU, microprocessor, multi-core processor, etc.) or a specialized processor (e.g., ASIC, DSP, FPGA, etc.).

[0026] Fig. Figure 2 is a block diagram that generally represents a signal processing stage of the audio device. Fig. Figure 1 represents the inputs to the signal processing stage 130. These inputs are time-domain audio samples 201 (TD samples). The time-domain samples 201 can be obtained by transforming the physical sound wave pressure into an equivalent analog signal representation (voltage or current) by a transducer (microphone), followed by an analog-to-digital converter (ADC) that converts the analog signal into digital audio samples. This digitized time-domain signal is then converted into a frequency-domain signal by the signal processing stage. The frequency-domain signal can be characterized by a variety of frequency bands 220 (i.e., frequency sub-bands, sub-bands, bands, etc.).In one implementation, the signal conditioning stage uses a weighted overlap-add (WOLA) filter bank, as disclosed, for example, in US patent US 6,236,731 B1 entitled "Filterbank Structure and Method for Filtering and Separating an Information Signal into Different Bands, Particularly for Audio Signal in Hearing Aids." The WOLA filter band used can encompass a short-time window (frame) length of R samples and N sub-frequency bands to transform the time-domain samples into their equivalent complex data representation in the sub-band frequency domain.

[0027] As in Fig. As shown in Figure 2, the signal conditioning stage 130 outputs a multitude of frequency subbands. Each non-overlapping subband represents frequency components of the audio signal within a range (e.g., + / - 125 Hz) of frequencies around a center frequency. For example, a first frequency band (i.e., BAND_0) may be centered at zero (DC) frequency and include frequencies in the range of approximately 0 to approximately 125 Hz; a second frequency band (i.e., BAND_1) may be centered at 250 Hz and include frequencies in the range of approximately 125 Hz to approximately 375 Hz; and so on for a number (N) of frequency bands.

[0028] The frequency bands 220 (i.e. BAND_0, BAND_1, etc.) can be processed to modify the audio signal 111 received at the audio device 100.

[0029] For example, the audio signal modifier stage 150 (see Fig. 1) Apply processing algorithms to the frequency bands to amplify the audio signal. Accordingly, the Audio Signal Modifier Stage 150 can be configured for noise reduction and / or speech / tone enhancement. The Audio Signal Modifier Stage 150 can also receive signals from one or more classifiers indicating the presence (or absence) of a specific audio signal (e.g., a tone), a specific audio type (e.g., speech, music), and / or a specific audio state (e.g., background type). These received signals can modify how the Audio Signal Modifier Stage 150 is configured for noise reduction and / or speech / tone enhancement.

[0030] As in Fig. As shown in Figure 1, a signal indicating the presence (or absence) of music can be received by the audio signal modifier stage 150 from a music classifier 140. This signal can cause the audio signal modifier stage 150 to apply one or more additional algorithms, eliminate one or more algorithms, and / or modify one or more algorithms it uses to process the received audio signal. For example, while music is being detected, a noise suppression level (i.e., attenuation level) can be reduced so that the music (e.g., a music signal) is not degraded by attenuation. In another example, a feedback suppressor can be controlled for carryover (e.g., false feedback detection), adjustment, and gain while music is being detected so that tones in the music are not suppressed.In yet another example, the bandwidth of the audio signal modifier stage 150 can be increased while music is being captured to improve the quality of the music, and then reduced while no music is being captured to save energy.

[0031] The music classifier is configured to receive the frequency bands 220 from the signal conditioning stage 130 and output a signal indicating the presence or absence of music. For example, the signal can include a first level (e.g., a logic high voltage) indicating the presence of music and a second level (e.g., a logic low voltage) indicating the absence of music. The music classifier 140 can be configured to continuously receive the bands and continuously output the signal, so that a change in the signal level correlates in time with the moment music begins or ends. As in Fig. As shown in Figure 1, the music classifier 140 can include a feature detection and tracking unit 200 and a combination and music detection unit 300.

[0032] Fig. Figure 3 is a block diagram that generally represents a feature capture and tracking unit of the music classifier. Fig. Figure 1 illustrates this. The feature capture and tracking unit comprises a plurality of decision-making units (i.e., modules, units, etc.). Each decision-making unit in the plurality is configured to capture and / or track a property (i.e., a feature) associated with the music. Because each unit is directed toward a single property, the algorithmic complexity required for each unit to produce an output (or outputs) is limited. Accordingly, each unit can require fewer bar cycles to determine an output than would be required to determine all of the musical properties using a single classifier. Additionally, the decision-making units can operate in parallel and provide their results together (e.g., simultaneously).Thus, the modular approach can consume less power to work in real time (as perceived by the user) than other approaches and is therefore well suited for hearing aids.

[0033] Each decision-making unit of the music classifier's feature capture and tracking unit can receive one or more (e.g., all) of the tapes from the signal processor. Each decision-making unit is configured to produce at least one output corresponding to a determination about a particular musical feature. The output of a given unit can be a two-level (e.g., binary) value (i.e., feature rating) indicating a yes or no answer (i.e., a right or a wrong answer) to the question, "Is the feature being captured at this time?" If a musical feature has a multitude of components (e.g., tones), a given unit can produce a multitude of outputs. In this case, each of the multitude of outputs can each represent a capture decision (e.g.,a feature rating equal to a logical 1 or a logical 0) with respect to one of the multitude of components. If a particular musical property has a temporal (i.e., time-varying) aspect, the output of a particular unit can correspond to the presence or absence of the musical property within a specific time window. In other words, the output of the particular unit tracks the musical properties with the temporal aspect.

[0034] Some possible musical characteristics that can be detected and / or tracked are a beat, a pitch (or pitches), and modulation activity. While each of these characteristics alone may be insufficient to accurately determine whether an audio signal contains music, combining them can increase the accuracy of the determination. For example, determining that an audio signal has one or more pitches (i.e., tonality) may be insufficient to determine that it is music, since a pure (i.e., time-constant) pitch may be included in an audio signal (i.e., exist within it) without being music. Determining that the audio signal also exhibits high modulation activity can help determine that the identified pitches are likely music (and not a pure pitch from another source). Furthermore, determining that the audio signal has a beat would strongly indicate that the audio signal contains music.Accordingly, the feature detection and tracking unit 200 of the music classifier 140 comprises a bar detection unit 210, a tone detection unit 240 and a modulation activity tracking unit 270.

[0035] Fig. Figure 4A is a block diagram that generally represents a clock acquisition unit of the feature acquisition and tracking unit of the music classifier according to a first possible implementation. The first possible implementation of the clock acquisition unit receives only the first subband (i.e., frequency band) (BAND_0) from signal conditioning 130, since a clock frequency is most likely to be found within the range of frequencies (e.g., 0 to 125 Hz) of this band. First, an instantaneous subband (BAND_0) energy calculation 212 is performed as: E0[n]=X2[n,0] where n is the current frame number, X[n,0] is the real BAND_0 data, and E0[n] is the instantaneous BAND_0 energy for the current frame. When a WOLA filter bank of signal conditioning stage 130 is configured to be in even stack mode, the imaginary part of BAND_0 (which would otherwise be 0 for any real input) is filled with a (real) Nyquist band value. Therefore, in even stack mode, E0[n] is calculated as: E0[n]=real{X[n,0]}2 E0[n] is then low-pass filtered before decimation 214 to reduce aliasing. One of the simplest and most power-efficient low-pass filters 214 that can be used is the first-order exponential smoothing filter: E0LFP[n]=αbd×E0LPF[n−1]+(1−αbd)×E0[n] where α bd the smoothing coefficient and E 0LFP [n] is the low-pass filtered BAND_0 energy. Next, E 0LFP [n] decimated by a factor of M 216, which is E b[m] generated, where m is the frame number at the decimated rate: FsR×M, where R is the number of samples in each frame n. At this decimated rate, the search for a possible clock signal occurs at each m = N. b carried out, wherein N b The length of the clock acquisition observation period is [missing information]. Screening at the reduced (i.e., decimated) rate can save energy consumption by reducing the number of samples to be processed within a given duration. The screening can be performed in various ways. An effective and computationally efficient method is the use of normalized autocorrelation. The autocorrelation coefficients can be determined as: ab[m,τ]=∑i=0NbEb[m−i]Eb[m−i+τ]∑i=0NbEb[m−i]2 where τ is the amount of delay at the decimated frame rate and a b[m,τ] are the normalized autocorrelation coefficients at the decimated frame number m and the delay value τ.

[0036] A beat detection (BD) decision 220 is then made. To decide that a beat is present, a b [m,τ] is evaluated over a range of τ delays and a search is then performed for the first sufficiently high local maximum of a b[m,τ] according to an assigned threshold. The sufficiently high criterion can provide a sufficiently strong correlation for the result to be considered a clock, with the associated delay value τ determining the clock period. If no local maximum is found, or if no local maximum is deemed sufficiently strong, the probability of a clock being present is considered low. While finding one instance that meets the criteria might be sufficient for clock detection, multiple results with the same delay value over several N increase the probability of a clock being present. b -Intervals significantly increase the probability. As soon as a clock cycle is detected, the detection status flag BD[m] is set. bd ] set to 1, where m bd the clock capture frame number at the rate FsR×M×Nb If no beat is detected, this is indicated by the detection status flag BD [m bd ] set to 0.

[0037] Determining the current tempo value is not explicitly required for time signature detection. However, if the tempo is required, the time signature detection unit can include a tempo determination that uses a relationship between τ and the tempo in bars per minute to: BPM=Fs×60R×M×τ

[0038] Since typical musical tempos range between 40 and 200 bpm, a b [m,τ] are evaluated only over the τ values ​​that correspond to this range, thus avoiding unnecessary calculations and minimizing the computational load. Consequently,-a b [τ] evaluated only in integer intervals between: τ=0.3×FsR×M and τ=1.5×FsR×M

[0039] The parameters R, α bd , N bThe bandwidth of the filter bank and the sharpness of the filter bank's sub-band filter are all correlated, and independent values ​​cannot be proposed. Nevertheless, the choice of parameter value directly affects the number of calculations and the efficiency of the algorithm.

[0040] For example, higher N b Higher M-values ​​yield more accurate results. Low M-values ​​may be insufficient to extract the clock signature, while high M-values ​​can lead to measurement aliasing that compromises clock capture. The choice of α bd is also with R, F s and linked to the filter bank properties, and an incorrectly set value can produce the same result as an incorrectly set M.

[0041] Fig. Figure 4B is a block diagram that generally represents a clock acquisition unit of the feature acquisition and tracking unit of the music classifier according to a second possible implementation. The second possible implementation of the band acquisition unit receives all sub-bands (BAND_0, BAND_1,..., BAND_N) from the signal conditioning unit 130. Each frequency band is low-pass filtered 214 and decimated 216 as in the previous implementation. Additionally, for each band, over the observation periods N bA variety of features (e.g., values ​​for energy mean, energy standard deviation, energy maximum, energy kurtosis, energy skewness, and / or energy cross-correlation) are extracted (i.e., determined, calculated, etc.) and fed as a feature set to a neural network 225. The neural network 225 can be a deep (i.e., multilayer) neural network with a single neural output, depending on the clock detection (BD) decision. The switches .S0, S1,..., S NSwitches can be used to control which tapes are used in the beat acquisition analysis. For example, some switches can be opened to remove one or more tapes that are thought to contain limited useful information. For instance, BAND_0 is thought to contain useful information pertaining to a beat and can therefore be included in beat acquisition (i.e., always included) by closing switch S0. Conversely, one or more higher-order tapes can be excluded from subsequent calculations (i.e., by opening their respective switches) because they may contain different information pertaining to a beat. In other words, while BAND_0 can be used to capture a beat, one or more of the other tapes (e.g., BAND_1 ... BAND_N) can be used to differentiate the captured beat between a musical bar and other beat-like tones (i.e.,(tapping, rattling, etc.) to further differentiate. The additional processing (i.e., power consumption) associated with each additional band can be balanced against the need for further clock detection discrimination, depending on the specific application. One advantage of this is... Fig. The advantage of the clock capture implementation shown in 4B is that it is adaptable to extract features from different bands as needed.

[0042] In one possible implementation, the multitude of extracted 222 features (e.g., for the selected bands) could include an energy average for the band. For example, a BAND_0 energy average (E b_µ ) are calculated as: Eb_μ[m]=1Nb∑i=0Nb−1Eb[m−i], where N b the observation period is (e.g., number of previous frames) and m is the current frame number.

[0043] In one possible implementation, the multitude of extracted 222 features (e.g., for the selected bands) could include an energy standard deviation for the band. For example, a BAND_0 energy standard deviation (E) could be b_σ )) are calculated as: Eb_σ[m]=∑i=0Nb−1(Eb[m−i]−Eb_μ[m])2Nb

[0044] In one possible implementation, the multitude of extracted 222 features (e.g., for the selected bands) can include an energy medium for the band. For example, a BAND_0 energy maximum (E) could b_max ) are calculated as: Eb_max[m]=max(Eb[m−i]|i=0i=Nb−1)

[0045] In one possible implementation, the multitude of extracted features 222 (e.g., for the selected bands) can include an energy kurtosis for the band. For example, a BAND_0 energy kurtosis (E b_k )) are calculated as: Eb_k[m]=1Nb∑i=0Nb−1(Eb[m−i]−Eb_μ[m]Eb_σ)4

[0046] In one possible implementation, the multitude of extracted 222 features (e.g., for the selected bands) could contain an energy skew for the band. For example, a BAND_0 energy skew (E b_s ) are calculated as: Eb_s[m]=1Nb∑i=0Nb−1(Eb[m−i]−Eb_μ[m]Eb_σ[m])3

[0047] In one possible implementation, the multitude of extracted 222 features (e.g., for the selected bands) can include an energy cross-correlation vector for the band. For example, a BAND_0 energy cross-correlation vector (E) could be b_xcor ) are calculated as: E¯b_xcor[m]=[ab[m,τ40],ab[m,τ40−1],…,ab[m,τ200+1],ab[m,τ200]] where τ is the correlation lag (i.e., the delay). The delays in the cross-correlation vector can be calculated as: τ200=round(0.3×FsR×M) and τ40=round(1.5×FsR×M)

[0048] While the present disclosure is not limited to the set of extracted features described above, these features can, in a possible implementation, form a feature set that a neural BD network 225 can use to determine a clock. One advantage of the features in this feature set is that they do not require computationally intensive mathematical calculations, thus saving processing power. Additionally, the calculations share common elements (e.g., mean, standard deviation, etc.), so the calculations of the shared elements only need to be performed once by the feature set, further saving processing power.

[0049] The neural BD network 225 can be implemented as a long short-term memory (LSTM) neural network. In this implementation, the entire cross-correlation vector (i.e., E̅) can be b_xcor[m]) by the neural network to reach a BD decision. In another possible implementation, the neural BD network 225 can be implemented as a forward neural network that uses a single max value of the cross-correlation vector, namely E max_xcor [m] to reach a BD decision. The implemented neural BD network of a specific type can be based on a balance between performance and power efficiency. For clock detection, the forward neural network can exhibit better performance and improved power efficiency.

[0050] Fig. Figure 5 is a block diagram that generally represents a tone detection unit 240 of the feature detection and tracking unit 200 of the music classifier 140 according to one possible implementation. The inputs to the tone detection unit 240 are the complex subband data from the signal state stage. While all N bands can be used to detect tonality, experiments have shown that subbands above 4 kHz may not contain enough information to justify the additional computations unless power efficiency is not a concern. Thus, for 0 < k < N TN , TN, where N TN The total number of subbands, in order to search for the presence of tonality, is calculated by dividing the instantaneous energy 510 of the complex subband data for each band as such: Einst[n,k]=|X[n,k]|2

[0051] Next, the band energy data are converted to log2. While a highly accurate log2 operation can be used, if the operation is considered too expensive, one that would approximate the results within fractions of dB may suffice, provided the approximation is relatively linear and monotonically increasing in its error. One possible simplification is the linear approximation, given as: L=E+2mr

[0052] where E is the exponent of the input value and m r The rest is... The approximation L can then be determined using a leading bit detector, two shift operations, and one add operation—instructions commonly found on most microprocessors. The log2 estimate of the instantaneous energy, called E inst_log[n,k], d is then processed by a low-pass filter 514 to remove interference from adjacent bands and concentrate on the frequency of the middle band in band k: Epre_diff[n,k]=αpre×Epre_diff[n−1,k]+(1−αpre)×Einst_log[n,k] where α pre the effective cutoff frequency coefficient and the resulting output divided by E pre_diff [n,k] or the predifferentiation filter energy. Next, a first-order differentiation 516 takes place in the form of a single difference over the current and previous frames of the R sampling: Δmag[n,k]=Epre_diff[n,k]−Epre_diff[n−1,k] and the absolute value of Δ mag is taken. The resulting output |Δ mag [n, k]| is then passed through a smoothing filter 518 to obtain an averaged |Δ mag [n,k]| over multiple time frames: Δmag_avg[n,k]=αpost×Δmag_avg[n−1,k]+(1−αpost)×|Δmag[n,k]| where αpost the exponential smoothing coefficient and the resulting output Δ mag_avg [n,k] is a pseudovariance measurement of the energy in band k and frame n in the logarithmic domain. Finally, two conditions are tested to decide (i.e., to determine) whether or not tonality is present: Δ mag_avg [n,k] is tested against a threshold below which the signal is considered to have a sufficiently low variance to be tonal, and E pre_diff [n,k] is checked against a threshold value to verify that the observed tonal component contains sufficient energy in the subband: TN[n,k]=(Δmag_avg[n,k]<TonalityTh[k])&&(Epre_diff[n,k]> SBMagTh[k]) where TN[n,k] contains the tonality presence status in band k and frame n at any given time. In other words, the outputs TD_0, TD_1,...TD_N can correspond to the probability that a tone is present within the band.

[0053] A common signal that is not music but contains some tonality, exhibits similar (to some music genres) temporal modulation properties, and possesses similar (to some music genres) spectral shapes to music is speech. Since it is difficult to robustly distinguish speech from music based on modulation patterns and spectral differences, the tonality level becomes the critical difference. The tonality threshold Th [k] must therefore be carefully selected so that it triggers only with music and not with speech. Since the value of tonality Th [k] from the pre- and post-differentiation filter set, namely the selected values ​​for α pre and α post, is dependent on, which itself depends on F s Since the threshold depends on the selected filter bank properties, no independent values ​​can be suggested. However, the optimal threshold can be obtained by optimizing a large database for a selected set of parameter values. While SBMag Th [k] also from the chosen α pre Since it depends on the -value, it is far less sensitive, as it merely serves to ensure that the detected tonality does not have too little energy to be insignificant.

[0054] Fig. Figure 6 is a block diagram that generally represents a modulation and activity tracking unit 270 of the feature acquisition and tracking unit 200 of the music classifier 140 according to one possible implementation. The input to the modulation activity tracking unit is the complex sub-band (i.e., band) data from the signal conditioning stage. All bands are combined (i.e., summed) for a broadband representation of the audio signal. The instantaneous broadband energy 610 E wb_inst [n] is calculated as: Ewb_inst[n]=∑k=0Nsb−1|X[n,k]|2 where X[n,k] is the complex WOLA (i.e., subband) with analysis data at frame n and band k. The broadband energy is then averaged over several frames using a smoothing filter 612: Ewb[n]=αw×Ewb[n−1]+(1−αw)×Ewb_inst[n] where α w the exponential smoothing coefficient and E wb[n] is the averaged broadband energy. Beyond this step, the modulation activity can be tracked to measure temporal modulation activity in different ways. 614 Some methods are more demanding, while others are computationally more efficient. The simplest and perhaps most computationally efficient method involves performing minimum and maximum tracking of the averaged broadband energy. For example, the global minimum value of the averaged energy could be captured every 5 seconds as the minimum energy estimate, and the global maximum value of the averaged energy could be captured every 20 ms as the maximum energy estimate. Subsequently, at the end of each 20 ms, the relative divergence between the min and max trackers is calculated and stored. r[mmod]=Max[mmod]Min[mmod] where m mod the frame number at the 20 ms interval rate, Max[m mod] the current estimate of the maximum value of broadband energy, Min[m mod ] the current (last updated) estimate of the minimum value of broadband energy and r[m mod ] the divergence ratio. The divergence ratio is then compared to a threshold value to determine a modulation pattern 616: LM[mmod]=(r[mmod] <Divergenzth)

[0055] The divergence value can take on a wide range. A low, medium, or high range would indicate an event that could be music, speech, or noise. Since the variance of the broadband energy of a pure tone is significantly low, an extremely low divergence value would indicate either a pure tone (of a volume level) or a non-pure tone signal at an extremely low level, which would most likely be too low to be considered desirable. The distinctions between speech and music, and noise and music, are made by tonality measurements (by the tonality detection unit) and the clock presence status (by the clock detection unit), and the modulation pattern or divergence value does not add much value in this regard.However, since pure tones cannot be distinguished from music by tonality measurements, and if present, may satisfy the tonality condition for music, and since the absence of a bar detection does not necessarily imply a non-musical condition, there is an explicit need for an independent pure tone detector. Since, as discussed, the divergence value can be a good indicator of whether a pure tone is present or not, we use the modulation pattern tracking unit exclusively as a pure tone detector to distinguish pure tones from music when the tone detection unit 240 determines that tonality is present. Consequently, we set the divergence. th at a sufficiently small value below which either only a pure tone or an extremely low signal (which is of no interest) can be present. Consequently, LM[m mod] or the low modulation status flag effectively becomes a "pure tone" or "non-music" status flag for the rest of the system. The output (MA) of the modulation activity tracking unit 270 corresponds to a modulation activity level and can be used to prevent a tone from being classified as music.

[0056] Fig. Figure 7A is a block diagram that generally represents a combination and music capture unit 300 of the music classifier 140 according to a first possible implementation. In a node unit 310 of the combination and music capture unit 300, all individual outputs of the individual capture units (i.e., feature ratings) (i.e., BD, TD_1, TD_2, TD_N, MA) are received and a weighting is applied (β). B , β T0 , β T1 , β TN , β M), to obtain a weighted feature score for each. The results are combined 330 to formulate a music score (e.g., for a frame of audio data). The music score can be accumulated over an observation period, during which a large number of music scores are obtained for a large number of frames. Period statistics 340 can then be applied to the music scores. For example, the obtained music scores can be averaged. The results of the period statistics are compared to a threshold 350 to determine whether music is present or absent during the period. The combination and acquisition unit is also configured to apply hysteresis control 360 to the threshold output to prevent potential language classifications from fluctuating between observation periods.In other words, a current threshold decision can be based on one or more passable threshold decisions. After the 360 ​​hysteresis control is applied, a final speech classification decision (MUSIC / NO MUSIC) is provided or made available to other subsystems in the audio device.

[0057] The Combination and Music Acquisition Unit 300 can operate on asynchronously arriving inputs from the acquisition units (e.g., Beat Acquisition 210, Tone Acquisition 240, and Modulation Activity Tracking 270) if they operate in different internal decision-making (i.e., determination) intervals. The Combination and Music Acquisition Unit 300 also operates in a highly computationally efficient manner while maintaining accuracy. At this high level, several criteria must be met for music to be acquired. For example, a strong beat or tone must be present in the signal, and the tone must not be a pure tone or an extremely low-level signal.

[0058] Since the decisions occur at different rates, the base update rate is set to the shortest interval in the system, which is the rate at which the tonality detection unit 240 operates on each R sample (the n frames). The feature evaluations (i.e., decisions) are weighted and thus combined into a music evaluation (i.e., rating):

[0059] In each frame n: B[n]=BD[mbd] M[n]=LM[mmod] where B[n] is updated with the latest clock detection status and M[n] is updated with the latest modulation pattern status. Then, for every N MD Interval: Score=0Score=∑i=0NMD−1(max(0,βBB[n−i]+∑k=0NTN−1βTKTN[n−i,k]+βMM[n−i]))Music Detected=(Score>MusicScoreth) where N (MD) the music recording interval length in frames, β B the weighting factor in connection with clock signal acquisition, β Tkthe weighting factor in connection with tonality detection is and β M The weighting factor is related to pure tone capture. The β weighting factors can be determined based on training and / or usage and are usually factory set. The values ​​of the β weighting factors can depend on several factors, which are described below.

[0060] First, the values ​​of the β weighting factors can depend on the significance of the event. For example, a single tonality hit may not be as significant for an event compared to a single bar detection event.

[0061] Secondly, the values ​​of the β weighting factors can depend on the internal tuning of the detection unit and the overall confidence level. It is generally advantageous to allow a small percentage of failure at the lower-level decision-making stages and to use long-term averaging to correct for some of this. This avoids setting overly restrictive thresholds at the lower levels, which in turn increases the overall sensitivity of the algorithm. The higher the specificity of the detection unit (i.e., a lower misclassification rate), the more significant the decision should be, and therefore a higher weighting value must be chosen. Conversely, the lower the specificity of the detection unit (i.e., a higher misclassification rate), the less significant the decision should be, and therefore a lower weighting value must be chosen.

[0062] Third, the values ​​of the β-weighting factors can depend on the internal update rate of the capture unit compared to the base update rate. Even if B[n], TN[n,k], and M[n] are all combined at each frame nB[n], M[n], the same state pattern persists for many consecutive frames due to the fact that the clock capture unit and the modulation activity tracking units update their flags at a decimated rate. For example, if BD[m bd ] running on an update interval period of 20 ms and the base frame period is 0.5 milliseconds, generates B[n] for each actual BD [m bd]-Hour clock capture event 40 consecutive frames of hour clock capture events. Thus, the weighting factors must take into account the multirate nature of the updates. If, in the example above, the intended weighting factor for an hour clock capture event was chosen to be 2, then β3 should be set to 2200.5 = 0.05 to be assigned to account for the 0.5 repetition pattern.

[0063] Fourth, the values ​​of the β weighting factors can depend on the correlation relationship of the detection unit's decision regarding music. A positive β weighting factor is used for detection units that support the presence of music, and a negative β weighting factor is used for those that reject the presence of music. Therefore, the weighting factors β B and β Tk positive weightings, while β m holds a negated weighting value.

[0064] Fifth, the values ​​of the β weighting factors can depend on the architecture of the algorithm. Since M[n] must be included in the summation node as an AND operation rather than an OR operation, a significantly higher weighting for β may be required. m are chosen to set the outputs of B[n] and TN[n,k] to zero and to act as an AND operation.

[0065] Even in the presence of music, not every music detection period necessarily detects music. Therefore, it may be desirable to accumulate several periods of music detection decisions before declaring the music classification to avoid potential music detection state fluctuations. It may also be desirable to remain in the music state longer if we have been in that state for an extended period. Both goals can be achieved very efficiently using a music state tracking counter: `if MusicDetected MusicDetectedCounter = MusicDetectedCounter + 1; else MusicDetectedCounter = MusicDetectedCounter - 1; end MusicDetectedCounter = max(0, MusicDetectedCounter) MusicDetectedCounter = min(MAX_MUSIC_DETECTED_COUNT, MusicDetectedCounter)` where `MAX_MUSIC_DETECTED_COUNT` is the value at which the `MusicDetectedCounter` is capped.A threshold value is then assigned to the MusicDetectedCounter, beyond which the music classification is declared. MusicClassification=(MusicDetectedCounter≥MusicDetectedCounterth)

[0066] In a second possible implementation of the combination and acquisition unit 300 of the music classifier 140, the weighting application and the combination process can be replaced by a neural network. Fig. Figure 7B is a block diagram that generally represents a combination and music acquisition unit of the music classifier according to the second possible implementation. The second implementation may consume more power than the first implementation ( Fig. 7A). Accordingly, the first possible implementation could be used for applications with lower available power (or modalities), while the second possible implementation could be used for applications with higher available power (or modalities).

[0067] The output of Music Classifier 140 can be used in various ways, and the usage depends entirely on the application. A fairly common result of a music classification state is the retuning of parameters in the system to better suit a musical environment. For example, in a hearing aid, when music is detected, any existing noise reduction might be disabled or tuned down to avoid any unwanted artifacts for the music. In another example, a feedback suppressor, while music is detected, does not respond to the observed tonality in the input in the same way as it would if no music were being detected (i.e., the observed tonality is due to feedback). In some implementations, the output of Music Classifier 140 (i.e.,MUSIC / NO MUSIC) are shared with other classifiers and / or stages in the audio device to help the other classifiers and / or stages perform one or more functions.

[0068] Fig. Figure 8 is a hardware block diagram, generally showing an audio device 100 according to one possible implementation of the present disclosure. The audio device includes a processor (or processors) 820, which can be configured by software instructions to perform all or some of the functions described herein. Accordingly, the audio device 100 also includes a memory 830 (e.g., non-volatile, computer-readable memory) for storing the software instructions and the parameters for the music classifier (e.g., weights). The audio device 100 may further include an audio input 810, which may include the microphone and the digitizer (A / D) 120. The audio device may further include an audio output 840, which may include the digital-to-analog (D / A) converter 160 and a loudspeaker 170 (e.g., a ceramic loudspeaker, a bone conduction loudspeaker, etc.).The audio device may further include a user interface 860. The user interface may include hardware, switching logic, and / or software for receiving voice commands. Alternatively or additionally, the user interface may include controls (e.g., buttons, selectors, switches) that a user can configure to adjust parameters of the audio device. The audio device may further include a power interface 880 and a battery 870. The power interface 880 may receive and process power to charge the battery 870 or to operate the audio device (e.g., regulate it). The battery may be a rechargeable battery that receives power from the power interface and may be configured to provide power for the operation of the audio device. In some implementations, the audio device may communicate with one or more computing devices 890 (e.g.,The audio device may be connected to a smartphone or a network (e.g., a cellular network, computer network). For these implementations, the audio device may include a communication interface (i.e., a COMM interface) to provide analog or digital communications (e.g., WiFi, BLUETOOTH™). The audio device may be portable and may be physically small and shaped to fit in the ear canal. For example, the audio device may be implemented as a hearing aid for a user.

[0069] Fig. Figure 9 is a flowchart of a method for capturing music in an audio device according to a possible implementation of the present disclosure. The method can be performed by hardware and software of the audio device 100. For example, a (non-volatile) computer-readable medium (i.e., memory) containing computer-readable instructions (i.e., software) can be accessed by the processor 820 to configure the processor to capture all or part of the music stored in the audio device. Fig. 9. carries out the procedure shown.

[0070] The process begins by receiving an audio signal (e.g., through a microphone). Receiving may include digitizing the audio signal to create a digital audio stream. Receiving may also include dividing the digital audio stream into frames and buffering the frames for processing.

[0071] The procedure further includes obtaining 920 subband (i.e., band) information corresponding to the audio signal. Obtaining the band information may (in some implementations) involve applying a weighted overlap addition (WOLA) filter bank to the audio signal.

[0072] The procedure further includes applying the tape information to one or more decision-making units. The decision-making units may include a beat detection (BD) unit configured to determine the presence or absence of a beat in the audio signal. The decision-making units may also include a tone detection (TD) unit (i.e., tonality detection unit) configured to determine the presence or absence of one or more tones in the audio signal. The decision-making units may also include a modulation activity (MA) tracking unit configured to determine the level (i.e., degree) of modulation in the audio signal.

[0073] The procedure further involves combining the results (i.e., the status, the state) of each of the one or more decision-making units. This combining may involve applying a weight to each output of the one or more decision-making units and then summing the weighted values ​​to obtain a musical score. The combining can be understood as similar to the computation associated with a node in a neural network. Accordingly, in some (more complex) implementations, the combining may involve applying the output of the one or more decision-making units to a neural network (e.g., a deep neural network, a forward neural network).

[0074] The procedure further excludes determining whether or not there is music in the audio signal from the combined results of the decision-making units. This determination may involve accumulating music ratings from frames (e.g., for a period of time, for a number of frames) and then averaging the music ratings. It may also involve comparing the accumulated and averaged music rating to a threshold. For example, if the accumulated and averaged music rating is above the threshold, music is considered to be present in the audio signal, and if the accumulated and averaged music rating is below the threshold, music is considered to be absent from the audio signal.The determination may also include applying hysteresis control to the threshold comparison, so that a previous state of music / no music influences the determination of the current state to prevent states of music present / no music from fluctuating back and forth.

[0075] The procedure further includes modifying the audio signal based on whether or not it contains music. This modification may include setting a noise reduction level so that the music levels are not reduced as if there were noise. The modification may also include disabling a feedback suppressor so that tones in the music are not suppressed as if they were feedback. The modification may also include increasing a passband for the audio signal so that the music is not filtered.

[0076] The procedure further includes the transmission of the modified audio signal. The transmission may include converting a digital audio signal into an analog audio signal using a D / A converter. The transmission may also include coupling the audio signal to a loudspeaker.

[0077] The disclosure can be implemented as a music classifier for an audio device. The music classifier includes a signal conditioning unit configured to transform a digitized time-domain audio signal into a corresponding frequency-domain signal encompassing a plurality of frequency bands; a plurality of decision-making units operating in parallel, each configured to evaluate one or more of the plurality of frequency bands to determine a plurality of feature ratings, each feature rating corresponding to a property associated with music; and a combination and music detection unit configured to combine the plurality of feature ratings over a period of time to determine whether the audio signal contains music.

[0078] In some possible implementations, the clock detection unit includes a neural clock detection network, but in others, the clock detection unit may be configured to detect a repeating clock pattern in a first frequency band (i.e., the lowest of the multitude of frequency bands) based on a correlation.

[0079] In one possible implementation, the music classifier's combination and music detection unit is a neural network that receives the multitude of feature ratings and returns a decision about whether or not there is music (i.e., a signal).

[0080] The disclosure can also be implemented as a method for capturing music. The method comprises receiving an audio signal; digitizing the audio signal to obtain a digitized audio signal; transforming the digitized audio signal into a plurality of frequency bands; applying the plurality of frequency bands to a plurality of decision-making units operating in parallel; obtaining a feature score from each of the plurality of decision-making units, the feature score from each decision-making unit corresponding to a probability that a particular musical feature is included in the audio signal; and combining the feature scores to capture music in the audio signal.

[0081] In one possible implementation, the music detection procedure further includes multiplying the feature rating of each of the multitude of decision-making units by a respective weighting factor to obtain a weighted rating of each of the multitude of decision-making units; summing the weighted ratings from the multitude of decision-making units to obtain a music rating; accumulating music ratings across a multitude of frames of the audio signal; averaging the music ratings from the multitude of frames of the audio signal to obtain an average music rating; and comparing the average music rating to a threshold to detect music in the audio signal.

[0082] In another possible implementation, the music capture procedure further includes modifying the audio signal based on the music capture; and transmitting the audio signal.

[0083] The revelation can also be implemented as a hearing aid. The hearing aid includes a signal processing stage and a music classifier stage. The music classifier stage includes a feature detection and tracking unit and a combination and music detection unit.

[0084] In one possible implementation of the hearing aid, the device further includes an audio signal modifier stage coupled with the signal processing stage and the music classifier stage. The audio signal modifier stage is configured to process the multiple frequency bands differently when a music signal is received compared to when no music signal is received.

[0085] Typical embodiments were disclosed in the patent specification and / or the figures. The present disclosure is not limited to such exemplary embodiments. The use of the term "and / or" includes any and all combinations of one or more of the associated listed elements. The figures are schematic representations and are therefore not necessarily drawn to scale. Unless otherwise indicated, specific terms have been used in a general and descriptive sense and not for the purpose of limitation.

[0086] The disclosure describes a variety of possible detection features and combination methods for robust and energy-efficient music classification. For example, the disclosure describes a beat detector based on a neural network that can use a variety of possible features extracted from a selection of (decimated) frequency band information. When specific mathematics is disclosed (e.g., a variance calculation for a tonality measurement), it can be described as cost-effective (i.e., efficient) from the standpoint of processing power (e.g., cycles, energy). While these aspects and others as described herein have been illustrated, numerous modifications, substitutions, changes, and equivalents are now apparent to the person skilled in the art. It is therefore understood that the invention is defined in the claims.

Claims

[1] Music classifier for an audio device, the music classifier comprising: a signal processing unit configured to transform a digitized time-domain audio signal into a corresponding frequency-domain signal encompassing a variety of frequency bands; a multitude of decision-making units operating in parallel, each configured to evaluate one or more of the multitude of frequency bands in order to determine a multitude of feature ratings, each feature rating corresponding to a property associated with music; and a combination and music recording unit configured to to receive the multitude of feature ratings asynchronously from the decision-making units, and to combine the multitude of feature ratings over a period of time to determine whether the audio signal includes music, wherein the multitude of decision-making units includes a modulation activity tracking unit configured to detect broadband modulation based on a minimum averaged energy and a maximum averaged energy of a sum of the multitude of frequency bands. [2] Music classifier for an audio device according to claim 1, wherein the plurality of decision-making units includes a clock detection unit and wherein the clock detection unit is configured to select one or more frequency bands from the plurality of frequency bands, to extract a plurality of features from each selected frequency band, to input the plurality of features from each selected frequency band into a neural clock detection network and to detect a repeating clock pattern based on an output of the neural clock detection network. [3] Music classifier for an audio device according to claim 2, wherein the plurality of features extracted from each selected frequency band form a feature set that includes an energy mean, an energy standard deviation, an energy maximum, an energy kurtosis, an energy skewness and an energy cross-correlation vector. [4] Music classifier for an audio device according to claim 1, wherein the plurality of decision-making units includes a tone detection unit configured to detect a tone in one or more of the plurality of frequency bands based on an amount of energy and an energy variance in each of the plurality of frequency bands. [5] Music classifier for an audio device according to claim 1, wherein the combination and music detection unit is configured to apply a weighting to each feature rating to obtain weighted feature ratings, to sum the weighted feature ratings to obtain a music rating, to accumulate music ratings for a plurality of frames, to calculate an average of the music ratings for the plurality of frames, and to apply hysteresis control to an output of the threshold for music or no music. [6] Method for capturing music in an audio signal, the method comprising: Receiving an audio signal; Digitizing the audio signal to obtain a digitized audio signal; Transforming the digitized audio signal into a multitude of frequency bands; Applying the multitude of frequency bands to a multitude of decision-making units operating in parallel; asynchronously obtaining a feature score from each of the plurality of decision-making units, wherein the feature score from each decision-making unit corresponds to a probability that a particular musical feature is included in the audio signal; and Combining feature ratings to capture music in the audio signal, wherein the multitude of decision-making units includes a modulation activity tracking unit, and wherein: Receiving a feature rating from the modulation activity tracking unit includes the following: Capturing a broadband modulation based on a minimum averaged energy and a maximum averaged energy of a sum of the multitude of frequency bands. [7] Method for music acquisition according to claim 6, wherein the plurality of decision-making units includes a beat acquisition unit, and wherein: Obtaining a feature rating from the clock acquisition unit includes the following: Capturing, based on a neural network, a repeating clock pattern across a multitude of frequency bands. [8] Method for music acquisition according to claim 6, wherein the plurality of decision-making units includes a sound acquisition unit, and wherein: Obtaining a feature rating from the tone detection unit includes the following: Capturing a tone in one or more of the multitude of frequency bands based on an energy quantity and energy variance in each of the multitude of frequency bands. [9] Hearing aid, comprehensive: a signal processing stage configured to convert a digitized audio signal into a variety of frequency bands; and a music classifier coupled with the signal processing stage, wherein the music classifier includes the following: a feature detection and tracking unit that includes a plurality of decision-making units operating in parallel, each decision-making unit being configured to generate a feature score corresponding to a probability that a particular musical feature is included in the audio signal; and a combination and music recording unit configured to To receive feature ratings asynchronously from the decision-making units, and to combine the feature ratings over a time period to detect music in the audio signal, wherein the combination and music detection unit is configured to generate a first signal indicating music while music is being detected in the audio signal, and is configured to generate a second signal indicating no other music signal, wherein the multitude of decision-making units includes a modulation activity tracking unit configured to detect broadband modulation based on a minimum averaged energy and a maximum averaged energy of a sum of the multitude of frequency bands.

Citation Information

Patent Citations

  • Hearing aid system and a method of operating a hearing aid system

    US20170180875A1