Apparatus and method for classifying an audio signal

The apparatus and method improve MMVD diagnosis by using machine learning to accurately classify heart murmurs into healthy and pathological categories, addressing the inaccuracy of current methods and enabling timely treatment.

JP2025536884APending Publication Date: 2025-11-12BOEHRINGER INGELHEIM VETMEDICA GMBH +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
JP2025518620
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-30
Filing Date
2023-09-26
Publication Date
2025-11-12

AI Technical Summary

Technical Problem

Current methods for diagnosing myxomatous mitral valve disease (MMVD) in dogs, particularly stages B1 and B2, are inaccurate and require expert verification, leading to potential misdiagnosis and delayed treatment.

Method used

An apparatus and method utilizing machine learning-based classifiers to analyze audio signals, specifically heart murmurs, through a two-stage classification process that includes preprocessing, feature extraction, and multiple machine learning algorithms to accurately classify heart sounds into healthy and pathological categories, and further sub-classify pathological murmurs by severity.

Benefits of technology

Enhances the accuracy and reliability of diagnosing MMVD by automatically distinguishing healthy from pathological heart murmurs and identifying their severity levels, reducing the need for expert intervention and improving timely treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025536884000001_ABST
    Figure 2025536884000001_ABST
Patent Text Reader

Abstract

The present disclosure relates to an apparatus (10) for classifying at least one audio signal (20), comprising an input interface (12) configured to receive input information (22) of the audio signal (20), a first trained machine learning-based classifier (16) configured to map the input information (22) to one of first and second classes (24; 26) of audio signals, a second trained machine learning-based classifier (18), and an output interface configured to output information on which class the audio signal (20) belongs to, wherein the second trained machine learning-based classifier (18) is configured to map the input information (22) of the audio signal belonging to the first class (24) of audio signals to one of a plurality of third classes (28) of audio signals if the audio signal (20) belongs to the first class (24) of audio signals.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to an apparatus and corresponding method for classifying audio signals that can be used to determine undesirable effects such as damage to machinery or illness in mammals. [Background technology]

[0002] The audio or acoustic signal may make it possible to determine undesirable effects such as damage to machinery or illness in mammals, such as non-human mammals, particularly canines, more particularly dogs.

[0003] As used herein, non-human mammals refer in particular to companion animals or pets, and these terms are understood as synonyms herein. Pets or companion animals refer to domesticated animals kept for pleasure rather than utility, such as felines, e.g., cats, canines, e.g., dogs, and horses. In particular, pets refer to dogs.

[0004] Myxomatous mitral valve disease (MMVD, also known as endocarditis and degenerative or chronic valvular heart disease) causes the mitral valve leaflets to prolapse into the left atrium of the heart. Complications of myxomatous mitral valve disease include infective endocarditis, mitral valve regurgitation, sudden death, and stroke.

[0005] MMVD is the most common heart disease in dogs. The MMVD staging system describes four basic stages of heart disease and heart failure: Stage A, Stage B (including B1 and B2), Stage C, and Stage D. Approximately 4.3 million dogs worldwide suffer from stage B2 heart disease each year due to misdiagnosis or delayed diagnosis. Stage B2 refers to asymptomatic dogs with more advanced mitral valve regurgitation that is hemodynamically severe, long-standing, and causes radiographic and echocardiographic evidence of left atrial and left ventricular dilation, meeting the criteria for clinical trials to identify dogs that would clearly benefit from initiating drug therapy to delay the onset of heart failure. Signs associated with the progression of MMVD include coughing, increased respiratory rate, shortness of breath, lethargy, decreased exercise capacity, loss of appetite, and brief periods of loss of consciousness. Causes include arrhythmia, severe coughing, and left atrial laceration.

[0006] The prevalence increases with age, affecting approximately 10% of dogs aged 5-8 years, 25% of dogs aged 9-12 years, and 35% of dogs aged 13 years and older. Older dogs, primarily small breeds (<20 kg), are affected, including miniature poodles, miniature schnauzers, Yorkshire terriers, and dachshunds. Another susceptible breed is the Cavalier King Charles Spaniel, which exhibits the unique characteristic of often developing mitral valve endometriosis at a young age. Larger breeds are significantly less affected.

[0007] A heart murmur, a deviation from a healthy heart sound, is the first and most important criterion for diagnosing MMVD. Veterinarians can hear the murmur with a stethoscope before owners notice anything unusual in their pet. Therefore, the disease can sometimes be detected during routine checkups, such as vaccinations. However, it is difficult for general veterinarians to diagnose and requires training, experience, and expert verification. Accurate staging of heart sounds is essential for diagnosing MMVD. Stages are B1 (no cardiomegaly), B2 (cardiomegaly), and C (acute heart failure or history of heart failure). Starting at stage B2, patients can be effectively treated with cardiovascular medications (e.g., pimobendan). There is a correlation between the stage of the disease and the intensity of the murmur. The murmur is caused by turbulent blood flow from the left ventricle through damaged mitral valve leaflets as it passes from the left atrium to the left atrium. The "grade" of the murmur is determined by the volume of the murmur. Another way to grade a heart murmur is to -Quiet (less than heart sound) = Grade I and II - Moderate sound (similar to a heartbeat) = Grade III -Loud sounds (louder than the heartbeat) = Grade IV -Droning (very loud and can be heard when the stethoscope is removed from the chest) = Grade V and VI There is.

[0008] MMVD is usually classified as mild, moderate, or severe. A quiet mitral valve sound (Grade I or II) almost always indicates mild MMVD, but as the sound progresses, there is often no correlation between the degree of sound and the severity of mitral regurgitation. Therefore, like stage B1, it is not usually directly treated with medications (e.g., pimobendan). Physicians will almost certainly administer appropriate medications to patients with grade IV or higher. A moderate sound requires an echocardiogram to confirm the absence of dilation; only then can treatment be initiated.

[0009] Identifying a heart murmur is the first and most important criterion for diagnosis (of MMVD). The examining physician listens to the heart with a stethoscope. Depending on their level of experience and training in cardiology, the animal is referred to a cardiologist, who will likely re-examine the heart sounds. This is followed by a clinical examination that typically includes:

[0010] X-ray examination Heart size: The left atrial cardiac shadow enlarges first, followed later by the left ventricular cardiac shadow. Displacement of the left trunk bronchus. Another important function of radiography is the evaluation of the pulmonary veins and lung fields. Pulmonary venous congestion is an indication for treatment. Pulmonary edema usually reveals alveolar shadows in the hilar region. Pulmonary congestion: First, the pulmonary veins appear congested, which can later be diagnosed as pulmonary edema (fluid accumulation in the lungs).

[0011] Echocardiography Measurement of atrial and ventricular size can reliably detect dilation (stage B1 or B2). -Can measure the contractile ability of the myocardium. Additionally, color Doppler echocardiography can be used to quantify the degree of heart failure.

[0012] Radiographs should always be performed if a murmur is dominant (ACVIM Guidelines 2019). The distinction between stages B1, B2, and C is determined by radiographic / echocardiographic examination. Stage B1 refers to asymptomatic dogs with no radiographic or echocardiographic evidence of cardiac remodeling in response to MMVD. Changes are observed but not severe enough to require treatment. Stage B2 refers to asymptomatic dogs with advanced, hemodynamically severe mitral regurgitation that has progressed long enough for radiographic or echocardiographic evidence of left atrial or left ventricular dilation to be evident, fulfilling clinical trial criteria to identify dogs that would benefit from medical treatment to delay the onset of heart failure. Stage C refers to dogs with current or past clinical signs of heart failure due to MMVD.

[0013] There are important differences in the treatment of dogs with acute heart failure that require hospitalization and dogs with heart failure that can be treated at home. Summary of the Invention [Problem to be solved by the invention]

[0014] It has been found that improvements in the analysis of audio or acoustic signals can lead to significant improvements in injury or disease determination, e.g., in terms of accuracy and reliability. It is therefore an object of the present invention to improve the automatic identification of injury or disease based on the analysis of audio signals. [Means for solving the problem]

[0015] This object is addressed by the subject matter of the appended claims.

[0016] According to a first aspect of the present disclosure, an apparatus for classifying at least one audio signal is proposed. The apparatus comprises an input interface configured to receive input information of the audio signal. The input information of the audio signal can be the audio signal itself or other characteristics. The apparatus further comprises a trained first machine learning-based classifier configured to map the input information of the audio signal to one of first and second classes of audio signals. The apparatus further comprises a trained second machine learning-based classifier configured to map the input information of the audio signal belonging to the first class of audio signals to one of a plurality of third classes of audio signals if the audio signal belongs to the first class of audio signals. The apparatus further comprises an output interface configured to output information on which class the audio signal belongs to.

[0017] The proposed device can automatically classify an audio signal into one of several classes based on information in the audio signal, which can be used, for example, to automatically identify and classify mechanical damage or heart disease.

[0018] The first trained machine learning-based classifier uses all audio signals to determine whether they belong to the first class or the second class. The predicted audio signals of the second class are not used in the second trained machine learning-based classifier. Instead, only the predicted audio signals of the first class, i.e., only a subset of all audio signals initially used, are entered into the second trained machine learning-based classifier. By subdividing the classification into two subsequent separate classification stages, the computational complexity of the machine learning-based classifiers involved can be reduced, which may lead to better classification results.

[0019] In some embodiments, the audio signal comprises a periodic (or quasi-periodic) audio pattern, such as the periodic sound of machinery, a train passing over sleepers, or multiple cycles of a heart sound. Periodic heart sounds may include, for example, the heart murmur of a (non-human) mammal, such as a dog.

[0020] In some embodiments, the apparatus further comprises a pre-processor configured to pre-process the audio signal and generate the input information for the audio signal. For example, the pre-processor comprises one or more filters configured to filter the audio signal. The one or more filters may include a low-pass filter, a high-pass filter, and / or a band-pass filter to remove unwanted frequency components and / or noise from the audio signal.

[0021] In some embodiments, the pre-processor is configured to extract from the audio signal a number of features that characterize the audio signal. The extracted features may have lower dimensionality (lower data size) than the audio signal. The extracted features can be considered as input information of / for the audio signal. For example, the pre-processor can be configured to extract time-domain features and / or frequency-domain features that characterize the audio signal based on time-domain and / or frequency-domain analysis methods. Examples of such domain analysis methods include Fourier transform, in particular fast Fourier transform, power spectral density, or wavelet decomposition transform.

[0022] Examples of features that can be extracted include: maximum, mean, median, standard deviation, variance, skewness, kurtosis, mean absolute deviation, 25th quantile, 75th quantile, entropy, zero-crossing rate, crest factor, duration of the first and / or second peak in a pattern, duration between the first and second peak in a pattern, duration between the second peak in a first pattern and the first peak in a subsequent pattern, Mel-Frequency Cepstral Coefficients (MFCCs), pitch chroma, spectral flatness, spectral kurtosis, spectral skewness, spectral tilt, spectral entropy, fundamental frequency, bandwidth, spectral centroid, spectral flux, spectral roll-off, class information, severity information, position information, race information, weight information, additional information, and / or other parameters, or combinations thereof.

[0023] In some embodiments, the pre-processor is configured to divide the (filtered) audio signal into multiple time windows, each time window including at least one period of a periodic audio pattern (e.g., a heartbeat period), and extract from each time window features that characterize the audio signal of the time window. Audio signals, such as a series of heartbeats or a series of side sounds resulting from rotating machinery, may have periodicity. By knowing / determining this periodicity, the audio signal can be subdivided into multiple windows, each window including at least one of the repeating / periodic audio patterns. This allows each repeating audio pattern to be analyzed independently of the others, for example, by comparing this audio pattern with known audio patterns. Alternatively, repeating audio patterns can be analyzed with respect to other audio patterns that follow each audio pattern.

[0024] It should be noted that the repeated audio patterns may be substantially equal to each other, similar to each other, contain one or more peaks of similar shape (shape of the respective amplitudes plotted over time) and / or contain one or more peaks of similar shape (shape of the heights plotted over time) and of similar amplitude values ​​at respective times within the window length. According to an embodiment, the window lengths may be equal. For example, the window length may be determined based on the repetition frequency of the repeated patterns. According to another variant, the boundary between two audio patterns may be determined so as to determine the window length of each time window. This means that the respective window lengths of each window may be determined separately.

[0025] In some embodiments, the first class of audio signals may represent irregular audio signals, and the second class of audio signals may represent regular audio signals. Here, an "irregular audio signal" may be understood as an anomalous (abnormal) audio signal, i.e., an audio signal whose signal course deviates from a normal or expected signal course. Conversely, a "regular audio signal" may represent an audio signal having a signal course that corresponds to a normal or expected signal course. For example, the first class of audio signals may represent a pathological heart murmur, and the second class of audio signals may represent a healthy heart sound. In some embodiments, the first class of audio signals may represent myxomatous mitral valve disease (MMVD).

[0026] In some embodiments, the plurality of third classes of audio signals are associated with different irregularity levels of the audio signal. For example, the plurality of third classes of audio signals may be associated with different pathological heart murmur levels. For example, if the first class of audio signals indicates a pathological heart murmur, the plurality of third classes may be associated with different severity levels of the pathological heart murmur, such as mild, moderate, or loud / thrilling.

[0027] In some embodiments, the trained first machine learning-based classifier is configured to implement a trained first boosting algorithm. The trained second machine learning-based classifier is configured to implement a trained second boosting algorithm. In machine learning, boosting is an ensemble meta-algorithm primarily for reducing bias and also variance in supervised learning, and is a family of machine learning algorithms that transform weak learners into strong learners. Most boosting algorithms consist of iteratively training weak classifiers with respect to a distribution and adding them to a final strong classifier. Examples of boosting algorithms include XGBoost (eXtreme Gradient Boosting) or AdaBoost (Adaptive Boosting), a statistical classification meta-algorithm.

[0028] In some embodiments, the trained first and second machine learning-based classifiers are of the same type (same model or algorithm) but differ in different training signals and model parameters. The first machine learning-based classifier can be trained based on a first ground truth audio signal including a first class and a second class of audio signals to enable the first machine learning-based classifier to determine whether an audio signal belongs to the first class or the second class of audio signals. Which of the first ground truth audio signals is associated with which of the first and second classes of audio signals is known in advance. The second machine learning-based classifier can be trained based on a second ground truth audio signal including the first class of audio signals but not the second class to enable the second machine learning-based classifier to determine whether an audio signal belongs to one of a plurality of third classes. Which of the second ground truth audio signals is associated with which of the third classes of audio signals is known in advance. For example, the training signal for the first machine learning-based classifier can include known audio signals (features) of both pathological and healthy heart sounds. The training signal for the second machine learning based classifier may include only known audio signals (features of) pathological heart murmurs.

[0029] In some embodiments, the device may further include a trained third machine learning-based classifier configured to map information of an audio signal belonging to the first class and one of the plurality of third classes to one of a plurality of fourth classes of audio signals if the audio signal belongs to one of the third classes of audio signals and satisfies additional criteria. The additional criteria may be based on (or include) the age and / or breed of the mammal (e.g., dog) to which the audio signal belongs. The plurality of fourth classes may correspond to a staging system for MMVD. The staging system may include stage A, stage B (B1, B2), stage C, and stage D. In particular, the plurality of fourth classes may include stages B1 and B2 of MMVD.

[0030] In some embodiments, the trained first, second, and third machine learning-based classifiers are of the same type (e.g., boosting algorithm) but differ in their respective training signals and model parameters. The third machine learning-based classifier can be trained based on third ground truth audio signals that are not of the first class of audio signals but are of the second class and that satisfy additional criteria, and the third machine learning-based classifier can determine whether the audio signal belongs to one of a plurality of fourth classes. Which of the third ground truth audio signals is associated with which of the fourth classes of audio signals is known in advance. The additional criteria can be based on (or include) the age and / or breed of the mammal (e.g., dog) to which the audio signal belongs. That is, to train the third machine learning-based classifier, only third ground truth audio signals corresponding to a specific age and / or breed and that belong to the first class but not the second class can be used.

[0031] In some embodiments, the input interface and / or the output interface are configured as wireless interfaces, which may allow convenient transfer of the audio signals and / or results to or from the device. For example, the audio signals may be sent to an app running on a smartphone or other mobile device that implements the device for classifying audio signals.

[0032] According to a further aspect of the present disclosure, a method for classifying an audio signal is proposed, the method comprising: receiving input information of at least one audio signal; classifying, by a first machine learning based classifier, the audio signal into one of a first and a second class of audio signals based on input information of the audio signal; If the audio signal belongs to a first class of audio signals, classifying, by a second machine learning based classifier, the audio signal into one of a third plurality of classes of audio signals based on the input information of the audio signal; outputting information about which class the audio signal belongs to; Includes.

[0033] In some embodiments, the method further comprises, during the training phase: training a first machine learning based classifier with ground truth audio signals comprising a first class and a second class of audio signals, such that the first machine learning based classifier is capable of determining whether an audio signal belongs to the first class or the second class of audio signals; training a second machine learning based classifier with ground truth audio signals that include the second class of audio signals but not the first class of audio signals, such that the second machine learning based classifier is able to determine whether the audio signals belong to one of a plurality of third classes, wherein it is known in advance which of the ground truth audio signals are associated with which of the third classes of audio signals.

[0034] In some embodiments, the method includes classifying the audio signal by a third machine learning based classifier into one of a plurality of fourth classes of audio signals based on input information of the audio signal if the audio signal belongs to one of the third classes of audio signals and satisfies additional criteria.

[0035] The third machine learning-based classifier can be trained based on third ground truth audio signals that are of the first class but not the second class of audio signals and that satisfy additional criteria, such that the third machine learning-based classifier can determine whether the audio signal belongs to one of a plurality of fourth classes. Which of the third ground truth audio signals is associated with which of the fourth classes of audio signals is known in advance. The additional criteria can be based on (or include) the age and / or breed of the mammal (e.g., dog) to which the audio signal belongs. That is, only third ground truth audio signals corresponding to a particular age and / or breed and that belong to the first class but not the second class can be used to train the third machine learning-based classifier.

[0036] In some embodiments, the audio signals include periodic heart sounds, a first class of audio signals indicative of pathological heart murmurs, a second class of audio signals indicative of healthy heart sounds, and a plurality of third classes of audio signals associated with different severity levels of the pathological heart murmurs.

[0037] According to a further aspect of the present disclosure, there is proposed a computer program having a program code for performing the above method when the computer program is run on a computer, a processor or a programmable hardware element.

[0038] As indicated above, a potential application of embodiments of the present disclosure is the diagnosis of illness in non-human mammals, particularly animals such as dogs. Thus, according to embodiments, the audio signal may be a record of a heartbeat sequence of a dog or other animal or other non-human mammal, and / or a record of a heart murmur sequence of a dog or other animal or other non-human mammal.

[0039] Some examples of apparatus and / or methods are now described, by way of example only, with reference to the accompanying figures. [Brief explanation of the drawings]

[0040] [Figure 1] 1 shows a schematic block diagram of an apparatus for classifying at least one audio signal according to the present disclosure; [Figure 2] 1 shows an example of an audio signal corresponding to periodic heart sounds. [Figure 3] 1 shows a flowchart of a method for classifying at least one audio signal according to the present disclosure. [Figure 4] 1 illustrates a flowchart of a method for classifying at least one audio signal according to one embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION

[0041] Some embodiments will now be described in more detail with reference to the accompanying drawings. However, other possible embodiments are not limited to the features of these embodiments described in detail. Other embodiments may include feature modifications, as well as feature equivalents and substitutions. Furthermore, the terms used herein to describe particular embodiments should not be construed as limiting further possible embodiments. Throughout the description of the figures, identical or similar reference numerals indicate identical or similar elements and / or features, which may be implemented in the same or modified form while providing the same or similar function. Additionally, the thickness of lines, layers, and / or regions in the figures may be exaggerated for clarity.

[0042] Where two elements A and B are combined using "or", this is understood to disclose all possible combinations, i.e. A only, B only, and A and B, unless expressly stated otherwise in individual cases. As alternative expressions for the same combination, "at least one of A and B" or "A and / or B" can be used. This applies equally to combinations of more than two elements.

[0043] Where singular forms such as "a," "an," and "the" are used and the use of only a single element is not explicitly or implicitly defined as required, further embodiments may use multiple elements to implement the same functionality. Where a function is described below as being implemented using multiple elements, further embodiments may implement the same functionality using a single element or a single processing entity. Furthermore, it should be understood that the terms "comprise," "including," "comprises," and / or "comprising," when used, describe the presence of specified features, integers, steps, operations, processes, elements, components, and / or groups thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, processes, elements, components, and / or groups thereof.

[0044] 1 shows a schematic block diagram of an apparatus 10 for classifying an audio signal. The apparatus 10 may be a programmable hardware device including a memory and one or more processing units, such as a CPU and / or a GPU. The apparatus 10 may be a mobile phone, a tablet, a personal computer, etc. In other embodiments, the apparatus 10 may be implemented by a central server.

[0045] The device 10 includes an input interface 12 configured to receive at least one audio signal 20 or other characteristic information. The audio signal 20 may be a digital representation of an acoustic signal. For example, the audio signal 20 may be received in a waveform audio file format (wav file). The audio signal 20 may be associated with, for example, the heartbeat of a mammal (e.g., a dog) and thus may include a periodic or quasi-periodic audio pattern, such as multiple cycles of a periodic heart sound. A heart murmur is a distinctive heart sound produced when blood flows across a heart valve or blood vessel. FIG. 2 shows an example of an audio signal 20 corresponding to a periodic heart sound. For example, one heart sound cycle may be defined as the duration from a signal (peak) S1 to the next subsequent signal (peak) S1. Subsequent cycles of heart sounds may typically differ slightly from one another in the duration and / or amplitude of signal peaks S1, S2. Those skilled in the art having the benefit of this disclosure will understand that audio signal 20 may also be associated with other acoustic signals, such as, for example, acoustic signals of mechanical origin, without departing from the principles proposed herein.

[0046] The input interface 12 of the device 10 can be a wired or wireless interface, such as a universal serial bus (USB), a WiFi interface, a Bluetooth interface, an infrared interface, or a cellular communication interface. In this manner, the audio signal 20 (or other characteristic information) can be conveniently transferred from a data source to the device 10.

[0047] Optionally, the device 10 may comprise a pre-processor 14 configured to filter the audio signal 20 to remove or reduce undesired signal components, such as noise. Thus, the pre-processor 14 may comprise one or more digital filters, such as a low-pass filter, a band-pass filter, or a high-pass filter.

[0048] Additionally or alternatively, the optional pre-processor 14 can be configured to extract from the audio signal 20 a plurality of features 22 that characterize the audio signal 20. The pre-processor 14 can be configured to extract time-domain and / or frequency-domain features from the audio signal 20, which characterize or identify the audio signal 20. Examples of domain analysis methods include Fourier transforms, particularly fast Fourier transforms, power spectral density, or wavelet decomposition transforms. Feature extraction starts from an initial set of measured data (audio signal 20) and constructs derived values ​​(features) that are intended to be informative and non-redundant, facilitating subsequent learning and generalization steps and potentially improving human interpretation. Feature extraction is associated with dimensionality reduction. Thus, the dimensionality of the extracted plurality of features 22 is smaller than the dimensionality of the audio signal 20. The plurality of features 22 can also be considered a feature vector containing feature information of the audio signal.

[0049] The pre-processor 14 can be configured to automatically extract characteristic features of the audio signal 20, such as mammalian heartbeats. To this end, the audio signal 20 can be decomposed into multiple time windows, each containing a cardiac cycle. Accordingly, the feature extraction processor 14 can be configured to divide the audio signal 20 into multiple time windows. For example, determining the time windows can be based on an algorithm that finds repetitions within the audio signal 20. Each time window can contain at least one period of a periodic audio pattern, such as a cardiac cycle. The feature extraction processor 14 can be configured to extract from each time window a feature 22 that characterizes the audio signal of the time window.

[0050] The window length can be determined based on the duration of the audio signal 20 and the number of cycles of the periodic audio pattern. The calculation can be performed by simple division. Of course, according to further embodiments, the window length can be determined in a different way, for example by determining the duration of each cycle, for example, the interval between a systole (diastole) and a subsequent systole (diastole), and averaging these durations. According to further embodiments, the window length can change over time, for example if the periodicity of the pattern changes. This can happen, for example, if the heart rate decreases in the current situation.

[0051] Possible time domain features that can be extracted from the audio signal 20 or windows thereof are the mean of the audio pattern, the median of the audio pattern, the standard deviation of the audio pattern, the variance of the audio pattern or its variance relative to another pattern, the skewness of the audio pattern, the kurtosis of the audio pattern, the mean absolute deviation of the audio pattern, the 25th quantile of the audio pattern, the 75th quantile of the audio pattern, the entropy of the audio pattern, the zero crossing rate of the audio pattern, the quest factor, the duration of the first peak, the duration of another peak, the duration from the end of S1 to the start of the next S1, and the duration from the end of S2 to the start of the next S1. In particular, duration features become more meaningful when the audio signal 20 is divided into windows.

[0052] Possible frequency domain features that may be extracted from the audio signal 20 or a window thereof are Mel-frequency cepstral coefficients, pitch chroma, spectral flatness, spectral kurtosis, spectral skewness, spectral tilt, spectral entropy, dominant frequency, bandwidth, spectral centroid, spectral flux, and / or spectral roll-off.

[0053] It should be noted that the different feature types and the list of feature types are not limited to those mentioned. According to embodiments, feature extraction is performed primarily or fully automatically. In particular, windowing can be performed automatically.

[0054] The apparatus 10 further comprises a first trained machine learning-based classifier 16 configured to map information of the audio signal 20 to one of a first and second class of audio signals 24, 26. In one embodiment, the information of the audio signal corresponds to optionally extracted features 22, and in another embodiment, corresponds to the audio signal 20 itself. In other words, the first trained machine learning-based classifier 16 is configured to determine whether the audio signal 20 or its extracted features 22 is indicative of either the first class of audio signals 24 or the second class of audio signals 26.

[0055] In some embodiments, the first class of audio signals 24 may represent irregular audio signals, and the second class of audio signals 26 may represent regular audio signals. Here, an "irregular audio signal" may be understood as an abnormal audio signal, i.e., an audio signal whose signal course deviates from a normal or expected signal course. Conversely, a "regular audio signal" may represent an audio signal having a signal course that corresponds to a normal or expected signal course. In an example involving an audio signal 20 including heart sounds, the first class of audio signals 24 may represent a pathological heart murmur, and the second class of audio signals 26 may represent a healthy heart sound. Thus, the first class of audio signals 24 may represent myxomatous mitral valve disease (MMVD), and the second class of audio signals 26 may represent a healthy (non-human) mammal (e.g., a dog).

[0056] The trained first machine learning based classifier 16 can be based on, for example, a neural network, a random forest algorithm, or a boosting algorithm, in particular the XGboost algorithm.

[0057] The apparatus 10 further includes a trained second machine learning-based classifier 18 configured to map the extracted features 22 or the audio signals 20 belonging to the first class 24 of audio signals to one of a plurality of third classes 28A-C of audio signals. Thus, the trained second machine learning-based classifier 18 maps the extracted features 22 or the audio signals belonging to the first class 24 to either class 28-A, class 28-B, or class 28-C. Those skilled in the art having the benefit of this disclosure will understand that a different number (or amount) of third classes 28 are possible. The trained second machine learning-based classifier 18 is used only for audio signals 20 (or characteristic features thereof) belonging to the first class 24 of audio signals (e.g., pathological heart murmurs) and is not used for audio signals 20 (or characteristic features thereof) belonging to the second class 26 of audio signals (e.g., healthy heart sounds).

[0058] The plurality of third classes 28A-C of audio signals may be associated with different irregularity levels of the audio signal 20. That is, the plurality of third classes 28A-C of audio signals may indicate the degree to which the audio signal 20 deviates from a normal or expected audio signal (e.g., a healthy heart sound). For example, the plurality of third classes 28A-C of audio signals may be associated with different pathological heart murmur levels. For example, if the first class 24 of audio signals indicates a pathological heart murmur or MMVD, the plurality of third classes 28A-C may be associated with different severity levels of the pathological heart murmur, such as mild, moderate, or loud / thrilling. Thus, the output of the trained second machine learning based classifier 18 can indicate whether the audio signal 20 contains a mild pathological heart murmur (class 28-A, mild deviation from healthy heart sounds), a moderate pathological heart murmur (class 28-B, moderate deviation from healthy heart sounds), or a loud / thrilling pathological heart murmur (class 28-C, large deviation from healthy heart sounds).

[0059] The trained second machine learning-based classifier 18, like the trained first machine learning-based classifier 16, can be based on, for example, a neural network, a random forest algorithm, or a boosting algorithm, in particular the XGboost algorithm. In particular, the trained first and second machine learning-based classifiers 16, 18 can rely on the same type of machine learning algorithm or model. For example, the trained first and second machine learning-based classifiers 16, 18 can both rely on the same boosting algorithm and be programmed using the Python programming language.

[0060] However, the first and second machine learning-based classifiers 16, 18 may be trained using different training signals (ground truth data). During an initial training stage before the inference stage, the first machine learning-based classifier 16 may be trained with a ground truth audio signal that includes both a first class 24 (e.g., pathological heart murmurs) and a second class 26 (e.g., healthy heart sounds) of audio signals, and the first machine learning-based classifier 16 may determine whether the audio signal 20 belongs to the first class 24 or the second class 26 of audio signals. As is typical for ground truth data, it is known in advance which of the ground truth audio signals (or their characteristic information) for the first machine learning-based classifier 16 are associated with which of the first and second classes 24, 26 of audio signals. Alternatively, the second machine learning based classifier 18 can be trained with ground truth audio signals that include only the first class 24 (e.g., pathological heart murmurs) and no audio signals of the second class (e.g., healthy heart sounds), and the second machine learning based classifier 18 can determine whether the audio signal 20 belongs to one of a plurality of third classes 28A-C. As is typical for ground truth data, it is known in advance which of the ground truth audio signals (first class audio signals) for the second machine learning based classifier 18 are associated with which of the third classes 28A-C of audio signals.

[0061] Due to different training data, the trained first machine learning based classifier 16 is configured differently from the trained second machine learning based classifier 18. In particular, the trained first machine learning based classifier 16 and the trained second machine learning based classifier 18 may be different trained XGboost classifiers. The XGboost classifier is an example of a classifier based on a boosting algorithm.

[0062] Device 10 further includes an output interface (not shown) configured to output information regarding whether audio signal 20 belongs to class 24, 26, or 28A-C. The output interface may include a display and / or a wired or wireless interface, such as, for example, a universal serial bus (USB), a WiFi interface, a Bluetooth interface, an infrared interface, or a cellular communication interface. In this manner, the output information may be conveniently transferred from device 10 to a data recipient.

[0063] Those skilled in the art having the benefit of this disclosure will understand that the present apparatus may be used not only to automatically classify a single audio signal, but also to automatically classify multiple audio signals. In some embodiments, multiple audio signals may be classified sequentially or in parallel. Apparatus 10 may be implemented by a single programmable hardware device or by multiple programmable hardware devices. Apparatus 10 is therefore configured to execute a computer-implemented method 30 for classifying audio signals. A flowchart of computer-implemented method 30 is shown in FIG. 3.

[0064] The method 30 includes an act 32 of receiving at least one audio signal 20. The audio signal 20 may be associated with the heartbeat of a (non-human) mammal (e.g., a dog). Additionally, the method 30 may include an optional act 34 of (automatically) extracting from the audio signal 20 a plurality of features 22 characterizing the audio signal 20. As described above, the features 22 may include time-domain features and / or frequency-domain features characterizing the audio signal. The method 30 further includes an act 36 of classifying, by a first machine-learning-based classifier 16, the audio signal 20 into one of a first class and a second class of audio signals based on the audio signal 20 or the extracted features 22. The first class 24 may be associated with a pathological heart murmur, and the second class 26 may be associated with a healthy heart sound. Those skilled in the art having the benefit of this disclosure will understand that the first class may similarly be associated with other abnormal sounds (e.g., abnormal mechanical sounds), and the second class 26 may similarly be associated with other normal sounds (e.g., normal mechanical sounds). Only if the audio signal 20 belongs to the first class of audio signals 24, method 30 proceeds to act 38 of classifying, by the second machine learning-based classifier 18, the audio signal 20 into one of a plurality of third classes 28 of audio signals based on the extracted features 22 or the audio signal 20 itself. The third classes 28 may be associated with different severity levels of pathological heart murmurs. Those skilled in the art having the benefit of this disclosure will understand that the third classes 28 may similarly be associated with different severity levels of other abnormal sounds (e.g., abnormal mechanical sounds). Still further, method 30 includes act 39 of outputting information regarding which class the audio signal 20 belongs to.

[0065] Additionally, method 30 may further include classifying the audio signal into one of a plurality of fourth classes by a third machine learning-based classifier. This classification may be based on the audio signal or its features, and may be performed only if the audio signal belongs to one of the third classes of audio signals and meets additional criteria. Below, an embodiment with three sequential machine learning-based classifiers is described.

[0066] 4, there is shown a detailed flowchart of a computer-implemented method 40 for classifying an audio signal 20. The method 40 can be executed on a smartphone or tablet, for example in the form of an app.

[0067] Initially, the first machine learning-based classifier 16 (e.g., a boosting algorithm), the second machine learning-based classifier 18 (e.g., a boosting algorithm), and the third machine learning-based classifier 42 (e.g., a boosting algorithm) are provided with previously trained first model parameters 41. This may be in the form of one or more pickle files that can be used to save the machine learning models and serialize Python object structures. Once the first machine learning-based classifiers 16, 18, 42 are initialized, at least one audio file containing an audio signal 20 may be provided. The audio signal 20 may, for example, represent the heart sounds of a dog (or other mammal) recorded with a stethoscope. The one or more audio files 20 may be provided, for example, as ".wav files," and may be provided via a wireless interface. In addition to the audio files, information regarding the dog's age (an integer) and the dog's breed (a string) from which the heart sounds originated may be provided via the wireless interface.

[0068] The audio signal 20 extracted from the audio file may be provided to a pre-processor 14. The pre-processor 14 may include a band-pass filter 14-1, a cardiac cycle detection unit 14-2, and a feature extraction unit 14-3. The band-pass filter 14-1 may have a passband of, for example, 50-500 Hz. Furthermore, the first 0.5 s and the last 0.5 s of the audio signal 20 may be cut to remove unwanted acoustic signals. The cardiac cycle detection unit 14-2 may be configured to perform the above-mentioned windowing to obtain one or more time windows of the audio signal 20, where the time windows may correspond to at least one cardiac cycle of the heart sounds. The feature extraction unit 14-3 may be configured to extract a plurality of characteristic features 22 from one or more windows of the audio signal 20.

[0069] The characteristic features 22 can then be provided as input to a first machine learning-based classifier 16, which is configured to map the input to a first model output indicative of a first class 24 or a second class 26 of audio signals. That is, depending on the input features 22, the first machine learning-based classifier 16 predicts whether the audio signal 20 is indicative of a pathological heart murmur or a healthy heart sound. If the audio signal 20 is indicative of a healthy heart sound (class 26), this information is output and the method 40 ends or returns. Next, a new audio signal can be analyzed, for example.

[0070] On the other hand, if the audio signal 20 indicates a pathological heart murmur (class 24), the extracted features 22 of the audio signal 20 can be provided as input to a second machine learning-based classifier 18 configured to map the input to a second model output indicative of one of three classes 28-A, 28-B, and 28-C of audio signals. The three classes 28-A, 28-B, and 28-C indicate mild, moderate, and loud / thrilling pathological heart murmurs. That is, depending on the input features 22, the second machine learning-based classifier 18 will predict whether the audio signal 20 indicates a mild, moderate, or loud / thrilling pathological heart murmur. The respective information is output, and the method 40 can end or return.

[0071] If the dog's age is below a predetermined age threshold (here, 2 years), the pathological heart murmur can be classified as mild, moderate, and loud / thrilling congenital heart murmur. Congenital means that the dog is born with the condition. If the dog's age is above a predetermined age threshold (here, 2 years), it is classified as Mitral Valvular Disease (MMVD) or Dilated Cardiomyopathy (DCM). The respective information is output and method 40 can end or return.

[0072] If the breed is a small or medium breed, the pathological heart murmurs are classified as mild, moderate, and loud / thrilling MMVD murmurs. If the breed is a large breed, the pathological heart murmurs are classified as mild, moderate, and loud / thrilling DCM murmurs. The respective information is output and method 40 can end or return.

[0073] Audio signals 20 classified by a trained third machine learning-based classifier 42 (e.g., a boosting algorithm) as mild, moderate, or loud / thrilling MMVD murmurs are then classified into one of two MMVD stages of cardiac disease: B1 and B2. Stage B1 refers to asymptomatic dogs with no radiographic or echocardiographic evidence of cardiac remodeling in response to MMVD, as well as dogs with remodeling changes that are not severe enough to meet current clinical trial criteria used to determine the need for treatment initiation. Stage B2 refers to asymptomatic dogs with more advanced mitral regurgitation that is hemodynamically severe, long-lasting, and causes radiographic and echocardiographic findings of left atrial and left ventricular dilation, meeting the clinical trial criteria used to identify dogs that would clearly benefit from initiating drug treatment to delay the onset of heart failure.

[0074] If the audio signal 20 indicates one of the pathological heart murmurs (class 24), mild, moderate, and loud / thrilling pathological heart murmurs (one of three classes 28-A, 28-B, and 28-C), the dog is two years old or older, and the breed is small / medium, the extracted features 22 of the audio signal 20 can be provided as input to a third machine learning-based classifier 42 configured to map the input to a third model output indicating one of two classes 44-A, 44-B. The two classes 44-A, 44-B indicate mild, moderate, and loud / thrilling MMVD heart murmurs of stage B1 or stage B2. That is, depending on the input features 22, the second machine learning-based classifier 18 predicts whether the audio signal 20 indicates a mild, moderate, or loud / thrilling MMVD heart murmur of stage B1 or stage B2. The respective information is output and the method 40 can end or return.

[0075] It should be noted that the third machine learning-based classifier 42 is trained based on the third ground truth audio signals of the first class 24 (pathological heart murmurs) and not the second class, and meets the additional criteria of dog age > 2 and dog breed = small / medium, i.e., only third ground truth audio signals corresponding to a particular age and / or dog breed and belonging to the first class 24 but not the second class can be used to train the third machine learning-based classifier 42.

[0076] The present disclosure proposes a concept that adds intensity detection to sound detection. For this purpose, a multipart / cascade algorithm is proposed, and the trained model, here a feature matrix, can be loaded into the algorithm in the form of a pickle file. Each new sample (audio signal) to be tested undergoes the same steps as the trained model. The new sample / sound can be filtered from noise (environmental noise, ...). For this purpose, various time-frequency analyses (Fast Fourier Transform, Power Spectral Density, Wavelet Decomposition) can be performed. Heart sounds can then be detected using different methods, resulting in multiple windows, each containing one cardiac cycle. Various features 22 can be calculated for each cardiac cycle (window) in both the time and frequency domains. Conventionally, the same steps are performed with a pickle file containing multiple data. Finally, a new test data set can be fed to the first classifier 16, which uses the entire data to determine whether it is a pathological or healthy heart sound. The predicted healthy data is not used by the second classifier 18. Only the pathological data set, i.e., only a subset of the first data set, is passed to the second classifier 18. This classifier 18 then examines the loudness of the murmur and returns whether it is a mild, moderate, or loud murmur. With this information, recommendations can be made to the physician. Dogs under the age of 2 are more likely to have congenital murmurs. Small / medium breeds over 2 years of age are more likely to have MMVD. Larger dogs usually have DCM (dilated cardiomyopathy). Boosting algorithms such as XGBoost and AdaBoost are particularly suitable.

[0077] Aspects and features described in connection with a particular one of the above embodiments may be combined with one or more of the other embodiments, replacing the same or similar features of the other embodiments, or introducing the features in addition to the other embodiments.

[0078] Embodiments may further be or relate to a (computer) program comprising program code for performing one or more of the above methods when the program is run on a computer, processor, or other programmable hardware component. Accordingly, the steps, acts, or processes of different ones of the above methods may also be performed by a programmed computer, processor, or other programmable hardware component. Embodiments may also be directed to program storage devices, such as digital data storage media, that are machine-readable, processor-readable, or computer-readable and that encode and / or contain machine-executable, processor-executable, or computer-executable programs and instructions. The program storage device may include or be, for example, a digital storage device, a magnetic storage medium such as a magnetic disk or magnetic tape, a hard disk drive, or an optically readable digital data storage medium. Other embodiments may include a computer, processor, control unit, (Field) Programmable Logic Array ((F)PLA), (Field) Programmable Gate Array ((F)PGA), Graphics Processor Unit (GPU), Application Specific Integrated Circuit (ASIC), Integrated Circuit (IC), or System on a Chip (SoC) system programmed to perform the steps of the methods described above.

[0079] Furthermore, it is understood that the disclosure of multiple steps, processes, operations, or functions disclosed in this specification or claims should not be construed to imply that these operations are necessarily order dependent, unless explicitly stated in individual instances or required for technical reasons. Thus, the above description does not limit the execution of some steps or functions to a particular order. Furthermore, in further embodiments, a step, function, process, or operation may include and / or be divided into several sub-steps, sub-functions, sub-processes, or sub-operations.

[0080] When aspects are described in the context of an apparatus or a system, these aspects should also be understood as descriptions of a corresponding method. For example, a block, device, or functional aspect of an apparatus or system may correspond to a feature, such as a method step, of a corresponding method. Thus, aspects described in the context of a method should also be understood as descriptions of a corresponding block, element, property, or functional feature of a corresponding apparatus or system.

[0081] The following claims are incorporated into the detailed description of this specification, with each claim standing on its own as a separate embodiment. It should also be noted that, although a dependent claim may refer to a specific combination with one or more other claims, other embodiments may also include a combination of the dependent claim with the subject matter of any other dependent or independent claim. Such combinations are expressly suggested by this specification, unless it is expressly stated in a particular case that a specific combination is not intended. Furthermore, features of a claim should also be included in other independent claims, even if that claim is not directly defined as dependent on those other independent claims. [Explanation of symbols]

[0082] 10 equipment 12 Input Interface 16 Trained first machine learning based classifier 18 Trained second machine learning based classifier 20 Audio Signals 22 Input information 24 First Class of Audio Signals 26 Second Class of Audio Signals 28 Third Class of Audio Signals

Claims

1. An apparatus (10) for classifying at least one audio signal (20), comprising: an input interface (12) configured to receive input information (22) of said audio signal (20); a trained first machine learning based classifier (16) configured to map the input information (22) of the audio signal to one of first and second classes (24; 26) of audio signals; a trained second machine learning based classifier (18); and an output interface configured to output information about which class the audio signal (20) belongs to; Equipped with The apparatus (10) is configured to, when the audio signal (20) belongs to the first class (24) of audio signals, map the input information (22) of the audio signal belonging to the first class of audio signals (24) to one of a plurality of third classes (28) of audio signals.

2. The apparatus of claim 1 , wherein the audio signal comprises a plurality of cycles of heart sounds.

3. 3. The apparatus (10) of claim 2, wherein the first class of audio signals (24) is indicative of pathological heart murmurs and the second class of audio signals (26) is indicative of healthy heart sounds.

4. 4. The apparatus (10) of claim 2 or 3, wherein the plurality of third classes (28) of audio signals are associated with different pathological heart murmur levels.

5. 5. The apparatus (10) of claim 1, further comprising a pre-processor (14) configured to extract, as the input information (22) of the audio signal (20), from the audio signal (20) a plurality of features (22) that characterize the audio signal.

6. 5. The apparatus (10) of claim 4, wherein the pre-processor (14) is configured to extract time-domain and / or frequency-domain features that characterize the audio signal (20).

7. 7. The apparatus (10) of claim 1, wherein the first trained machine learning based classifier (16) is configured to implement a first trained boosting algorithm and / or the second trained machine learning based classifier (18) is configured to implement a second trained boosting algorithm.

8. The apparatus (10) according to any of the preceding claims, wherein the trained first and second machine learning based classifiers (16; 18) are of the same type and are different for different respective training signals.

9. The first machine learning-based classifier (16) is trained based on first ground truth audio signals including the first class of audio signals (24) and the second class of audio signals (26), and the first machine learning-based classifier (16) is adapted to determine whether the audio signal (20) belongs to the first class of audio signals or the second class of audio signals, and which of the first ground truth audio signals is associated with which of the first class of audio signals and which of the second class of audio signals.

9. The apparatus of claim 8, wherein: the first class of audio signals is known in advance; the second machine learning based classifier is trained based on second ground truth audio signals that include the first class of audio signals but not the second class of audio signals; and the second machine learning based classifier is configured to determine whether the audio signal belongs to one of a plurality of third classes; and it is known in advance which of the second ground truth audio signals are associated with which of the third classes of audio signals.

10. 10. The apparatus (10) of claim 1, further comprising a trained third machine learning-based classifier (42), the trained third machine learning-based classifier (42) configured to map the input information (22) of an audio signal that belongs to any one of the plurality of third classes (28) of audio signals and satisfies additional criteria to one of a plurality of fourth classes of audio signals.

11. The apparatus (10) of claim 10, wherein the trained first, second, and third machine learning based classifiers (16; 18; 42) are of the same type and are different for each different training signal.

12. 12. The apparatus (10) of claim 10 or 11, wherein the third machine learning based classifier (42) is trained based on third ground truth audio signals that are of the first class (24) of audio signals that are not of the second class of audio signals and that satisfy additional criteria, such that the third machine learning based classifier (42) is able to determine whether the audio signal (20) belongs to one of the plurality of fourth classes (28), and it is known in advance which of the third ground truth audio signals is associated with which of the fourth classes of audio signals.

13. The device (10) according to any of claims 10 to 12, wherein the additional criteria are based on the age and / or breed of the mammal to which the audio signal (20) belongs.

14. 1. A method for classifying at least one audio signal, comprising: receiving input information (22) of said audio signal (20); classifying the audio signal (20) into one of a first and a second class (24; 26) of audio signals based on the input information (22) by a first machine learning based classifier (16); If the audio signal (20) belongs to the first class (24) of audio signals, classifying, by a second machine learning based classifier (18), the audio signal (20) into one of a plurality of third classes (28) of audio signals based on the input information (22) of the audio signal; outputting information about which class the audio signal (20) belongs to; A method comprising:

15. During the training phase, training the first machine learning-based classifier (16) with ground truth audio signals including a first class of audio signals (24) and a second class of audio signals (26), such that the first machine learning-based classifier (16) is capable of determining whether the audio signal (20) belongs to the first class of audio signals or the second class of audio signals; training the second machine learning based classifier (18) with ground truth audio signals that include the first class of audio signals (24) but not the second class of audio signals, such that the second machine learning based classifier (18) is capable of determining whether the audio signals belong to one of a plurality of third classes (28), wherein it is known in advance which of the ground truth audio signals are associated with which of the third classes of audio signals; The method of claim 14 further comprising:

Citation Information

Patent Citations

  • Heart valve abnormality analysis method, system and device based on convolutional neural network

    CN111759345A

  • Apparatus and method for analyzing heart sound frequency

    JP2009240527A

  • Auscultatory cardiac sound signal processing method, auscultatory cardiac sound signal processing apparatus, and auscultatory cardiac sound signal processing program

    JP2014233598A

  • Classifier ensemble for detection of abnormal heart sounds

    JP2019531792A

  • Method and system for identifying a physiological or biological condition or disease in a subject

    JP2021534939A