Apparatus and method for classifying audio signal
By designing a device and method that utilizes machine learning-based classifiers, the problem of difficulty in automatically identifying and classifying audio signals in the prior art is solved, and the accurate classification and segmentation of categories such as healthy heart sounds and pathological heart murmurs is achieved, and the efficiency and accuracy of disease diagnosis are improved.
Patent Information
- Application Number
- CN202380067093.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-30
- Filing Date
- 2023-09-26
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art is difficult to effectively and automatically identify and classify audio signals, such as machine damage or mammalian diseases, such as myxomatous mitral valve disease (MMVD) in dogs.
A device and method are designed to automatically classify audio signals using a machine learning-based classifier. The device includes an input interface, a preprocessor and a number of machine learning-based classifiers that classify audio signals into different categories such as healthy heart sounds or pathological heart murmurs and further subdivides them into different severity levels by preprocessing and feature extraction of audio signals.
It realizes automatic classification of audio signals, improves the accuracy and efficiency of identification of diseases such as pathological heart murmurs, and reduces the dependence on professional knowledge.
Smart Images

Figure CN119947652A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to an apparatus and a corresponding method for classifying audio signals, which can be used to determine undesired effects, such as damage to a machine or illness in a mammal. Background Art
[0002] The audio signal or acoustic signal may enable determination of undesired effects such as damage to a machine or disease in a mammal, such as a non-human mammal, particularly a canine, more particularly a dog.
[0003] Non-human mammals herein refer in particular to companion animals or pets, which terms are to be understood as synonyms herein. Pets or companion animals refer to domesticated animals raised for entertainment rather than utility, for example, felines (such as cats), canines (such as dogs and horses). Specifically, herein, pets refer to dogs.
[0004] Myxomatous mitral valve disease (MMVD, also known as endocarditis and degenerative or chronic heart valve disease) is a cause of prolapse of (some) of the mitral valve leaflets into the left atrium of the heart. Complications of myxomatous mitral valve disease include infective endocarditis, mitral valve regurgitation, sudden death, and stroke.
[0005] MMVD is the most common heart disease in dogs. A staging system for MMVD describes four basic stages of heart disease and heart failure: Stage A, Stage B (including B1 and B2), Stage C, and Stage D. Each year, approximately 4.3 million dogs worldwide suffer from stage B2 heart disease that is misdiagnosed or diagnosed late. Stage B2 refers to asymptomatic dogs with more advanced mitral regurgitation that is severe and long-standing enough to cause radiographic and echocardiographic findings of left atrial and ventricular enlargement that meet clinical trial criteria for identifying dogs that would clearly benefit from initial medical treatment to delay the onset of heart failure. Symptoms of MMVD as it progresses are coughing, increased respiratory rate, rapid breathing, listlessness, poor performance, reluctance to eat, and transient loss of consciousness. Causes are irregular heartbeat or severe coughing or due to a tear in the left atrium.
[0006] The prevalence increases with age, affecting about 10% of all dogs 5 to 8 years old, about 25% of all dogs 9 to 12 years old and 35% of all dogs over 13 years old. It mainly affects older dogs of small breeds (<20kg), such as: Miniature Poodle, Miniature Schnauzer, Yorkshire Terrier, Dachshund. Another susceptible breed is the Cavalier King Charles Spaniel, which has the characteristic that it usually develops endocarditis at a young age. Large dogs are affected much less frequently.
[0007] As a deviation from healthy heart sounds, heart murmurs are the first and most important criterion for the diagnosis of MMVD. A veterinarian can hear heart murmurs with the help of a stethoscope, even before the owner himself notices any changes in his own pet. Therefore, this disease can be detected during routine examinations (such as vaccination examinations). However, this is difficult for ordinary veterinarians to diagnose and requires training, experience and expert verification. In order to diagnose MMVD, accurate staging of heart murmurs is a prerequisite. Stage B1 (no heart enlargement), stage B2 (with heart enlargement) and stage C (acute or previous heart attenuation) are distinguished. Starting from stage B2, a drug (e.g., Pimobendan) can be used as a cardiovascular drug for effectively treating patients. There is a correlation with the stage of the disease and the noise intensity. The murmur is caused by turbulent blood flow from the left ventricle through the damaged leaflet of the mitral valve to the left atrium. The loudness of the murmur is thereby determined by its "grade". An alternative method for grading murmurs provides:
[0008] Quiet (quieter than heart sounds) = Grade I and II
[0009] Moderate (as loud as heart sounds) = Grade III
[0010] Loud (louder than heart sounds) = Grade IV
[0011] ●Buzzing (very loud, heard when the stethoscope is removed from the chest) = Grades V and VI.
[0012] MMVD is usually classified as mild, moderate or severe. Although a quiet mitral sound (I or II grade) almost always indicates mild MMVD, once the sound rises from then on, there is usually no correlation between the degree of the sound and the degree of mitral regurgitation. Therefore, like stage B1, this is usually not directly treated with a drug (e.g., pimobendan). Almost certainly, the physician will give patients of grade IV and above appropriate drug therapy. In the case of moderate murmurs, cardiac ultrasound must be used to ensure that there is no increase. Treatment can only be given at that time.
[0013] Heart murmur identification is the first and most important criterion for diagnosis of (MMVD). The examining physician listens to the heart sounds with the help of a stethoscope. Depending on the level of experience and training in cardiology, the animal is referred to a cardiologist. The veterinary cardiologist will most likely re-examine the heart sounds. This is usually followed by further clinical examinations such as:
[0014] X-ray
[0015] o Heart size: There is an enlargement first in the area of the left atrium and subsequently also in the area of the left ventricle.
[0016] oDisplacement of the left trunk bronchus.
[0017] o Another important function of X-rays is the assessment of the pulmonary vasculature and lung fields. If the pulmonary veins are congested, this is an indication for treatment. If pulmonary edema is present, alveolar shadows may be visualized, usually in the hilar region.
[0018] o Pulmonary congestion: First the pulmonary veins become congested, and then pulmonary edema (fluid in the lungs) is diagnosed.
[0019] Echocardiography
[0020] o The size of the atria and ventricles can be measured, allowing any enlargement to be reliably detected (B1 or B2 stages).
[0021] oIt measures the ability of the heart muscle to contract.
[0022] o In addition, color Doppler echocardiography can be used to quantify the degree of dysfunction.
[0023] X-rays should always be performed when one murmur is dominant (ACVIM Guidelines 2019). The differences between stages B1, B2, and C are apparent on radiographs / echocardiograms. Stage B1 describes asymptomatic dogs with no radiologic or echocardiographic evidence of cardiac remodeling in response to their MMVD. Changes are observed, but they are not severe enough to require treatment. Stage B2 refers to asymptomatic dogs with advanced mitral regurgitation that is severe and long-standing enough to cause left atrial and ventricular enlargement, and these findings meet clinical trial criteria for identifying dogs that would clearly benefit from medical treatment to delay the onset of heart failure. Stage C refers to dogs with current or prior clinical signs of heart failure caused by MMVD.
[0024] There are important treatment differences between dogs with acute heart failure that require hospitalization and dogs whose heart failure can be treated at home.
[0025] It has been found that improvements in the analysis of an audio or acoustic signal can lead to significant improvements in the determination of damage or disease, for example with respect to accuracy and reliability.It is therefore an object of the present invention to improve the automatic identification of damage or disease based on an analysis of an audio signal. Summary of the invention
[0026] This object is solved by the object of the attached technical solution.
[0027] According to a first aspect of the present invention, a device for classifying at least one audio signal is provided. The device includes an input interface configured to receive input information of the audio signal. The input information of the audio signal may be the audio signal itself or other characteristics thereof. The device further includes a trained first machine learning-based classifier configured to map the input information of the audio signal to one of a first category and a second category of audio signals. The device further includes a trained second machine learning-based classifier configured to map the input information of the audio signal belonging to the first category of audio signals to one of a plurality of third categories of audio signals if the audio signal belongs to the first category of audio signals. The device further includes an output interface configured to output information about which category the audio signal belongs to.
[0028] The proposed device can automatically classify the audio signal into one of a plurality of categories based on the information of the audio signal. For example, the proposed device can be used to automatically identify and classify damage to a machine or heart disease.
[0029] The trained first machine learning based classifier uses all audio signals and checks whether they belong to the first or second class. The predicted second class of the audio signal is not used for the trained second machine learning based classifier. Instead, only the predicted first class of the audio signal (i.e., only a subset of the total audio signal used at the beginning) will enter the trained second machine learning based classifier. By subdividing the classification into two subsequent and distinct classification phases, the computational complexity of the involved machine learning based classifiers can be reduced. This can also lead to better classification results.
[0030] In some embodiments, the audio signal comprises a cyclic (or quasi-periodic) audio pattern, such as, for example, a periodic sound of a machine, a train passing over a railroad tie, or multiple cycles of a heart sound. For example, the periodic heart sound may include a heart murmur of a (non-human) mammal, such as a dog.
[0031] In some embodiments, the apparatus further comprises a preprocessor configured to preprocess the audio signal and generate input information of or about the audio signal. For example, the preprocessor comprises one or more filters configured to filter the audio signal. The one or more filters may include low-pass, high-pass and / or band-pass filters for removing undesirable frequency components and / or noise from the audio signal.
[0032] In some embodiments, the preprocessor is configured to extract a plurality of features that characterize the audio signal from the audio signal. The extracted features may have a dimension lower than the audio signal (lower data size). The extracted features may be considered as input information of / about the audio signal. For example, the preprocessor may be configured to extract time domain features and / or frequency domain features that characterize the audio signal based on time domain and / or frequency domain analysis methods. Examples of such domain analysis methods include Fourier transforms, in particular fast Fourier transforms, power spectral density or wavelet decomposition transforms.
[0033] Examples of extractable characteristic features are:
[0034] a maximum value, a mean, a median, a standard deviation, a variance, a skewness, a kurtosis, a mean absolute deviation, a 25th quantile, a 75th quantile, an entropy, a zero crossing rate, a crest factor, a duration of a first peak and / or a second peak within a pattern, a duration between a first peak and a second peak within a pattern, a duration between a second peak of a first pattern and a first peak of a subsequent pattern, a mel frequency cepstral coefficient (MFCC), a pitch chroma, a spectral flatness, a spectral kurtosis, a spectral skewness, a spectral slope, a spectral entropy, a dominant frequency, a bandwidth, a spectral centroid, a spectral flux, a spectral roll-off, a category information, a severity information, a location information, a race information, a weight information, additional information and / or other parameters or a combination thereof.
[0035] In some embodiments, the preprocessor is configured to divide the (filtered) audio signal into a plurality of time intervals, each time interval including at least one cycle of a periodic audio pattern (e.g., a heartbeat cycle), and extract features of the audio signal that characterize the time interval from each time interval. An audio signal, such as a series of heartbeats or a series of side sounds originating from a rotating machine, may have a periodicity. By knowing / determining this periodicity, the audio signal may be subdivided into a plurality of intervals such that each interval includes at least one of a repeating / periodic audio pattern. This may enable, for example, the analysis of each repeating audio pattern independently of one another by comparing this audio pattern with a known audio pattern. Alternatively, the repeating audio pattern may be analyzed relative to another audio pattern following the respective audio pattern.
[0036] It should be noted that the repeated audio patterns may be substantially equal to each other, similar to each other, include one or more peaks of a comparable shape (shape of respective amplitude plotted over time) and / or include one or more peaks of comparable shapes (shape of height plotted over time) and comparable amplitude values at respective time points within an interval length, etc. According to an embodiment, the interval lengths may be equal. For example, the interval lengths may be determined based on a frequency of repetition of the repeated pattern. According to another variant, a boundary between two audio patterns may be determined in order to determine the interval lengths of respective time intervals. This means that the interval lengths of each interval may be determined separately.
[0037] In some embodiments, the first category of audio signals represents an irregular audio signal, and the second category of audio signals represents a regular audio signal. Here, an "irregular audio signal" may be understood as an abnormal (abnormal) audio signal. That is, an audio signal having a signal process that deviates from a normal or expected signal process. Conversely, a "regular audio signal" may represent an audio signal having a signal process that corresponds to a normal or expected signal process. For example, the first category of audio signals may represent a pathological heart murmur, and the second category of audio signals may represent a healthy heart sound. In some embodiments, the first category of audio signals may indicate myxomatous mitral valve disease (MMVD).
[0038] In some embodiments, the plurality of third categories of the audio signal are associated with different levels of irregularity of the audio signal. For example, the plurality of third categories of the audio signal may be associated with different levels of pathological heart murmurs. If the first category of the audio signal indicates a pathological heart murmur, the plurality of third categories may be associated with different levels of severity of the pathological heart murmur, such as mild, moderate, or loud / thrilling, for example.
[0039] In some embodiments, the trained first machine-learning based classifier is configured to implement a trained first boosting algorithm. The trained second machine-learning based classifier is configured to implement a trained second boosting algorithm. In machine learning, boosting is mainly used to reduce the bias and variance in supervised learning as an ensemble meta-algorithm and a series of machine learning algorithms that convert weak learners into strong learners. Most boosting algorithms consist of repeatedly learning weak classifiers relative to a distribution and adding them to a final strong classifier. Examples of boosting algorithms are XGBoost (extreme gradient boosting) or AdaBoost (adaptive boosting), which are statistical classification meta-algorithms.
[0040] In some embodiments, the trained first and second machine learning based classifiers are of the same type (same model or algorithm) and are distinguished by different training signals and by different model parameters. The first machine learning based classifier can be trained based on a first ground truth audio signal including a first category and a second category of audio signals, so that the first machine learning based classifier can determine whether the audio signal belongs to the first category or the second category of audio signals. It is known in advance which of the first ground truth audio signal is associated with which of the first and second categories of audio signals. The second machine learning based classifier can be trained based on a second ground truth audio signal including the first category of audio signals but not the second category, so that the second machine learning based classifier can determine whether the audio signal belongs to one of a plurality of third categories. It is known in advance which of the second ground truth audio signal is associated with which of the third category of audio signals. For example, the training signal for the first machine learning based classifier may include known (features) of the audio signals of pathological and healthy heart sounds. The training signal for the second machine learning based classifier may only include known (features) of the audio signals of pathological heart murmurs.
[0041] In some embodiments, the apparatus may further include a trained third machine learning-based classifier configured to map information of an audio signal belonging to the first category and belonging to one of the plurality of third categories to one of the plurality of fourth categories of audio signals if the audio signal belongs to one of the third categories of audio signals and satisfies an additional criterion. The additional criterion may be based on (or include) an age and / or a breed of a mammal (e.g., a dog) to which the audio signal belongs. The plurality of fourth categories may correspond to a staging system for MMVD. The staging system may include stage A, stage B (B1, B2), stage C, and stage D. Specifically, the plurality of fourth categories may include stages B1 and B2 for MMVD.
[0042] In some embodiments, the trained first, second, and third second machine-learning-based classifiers are of the same type (e.g., boosting algorithm) and are distinguished by different respective training signals and model parameters. A third machine-learning-based classifier may be trained based on a third real data audio signal of the first category of audio signals but not of the second category and satisfying an additional criterion, so that the third machine-learning-based classifier is able to determine whether the audio signal belongs to one of a plurality of fourth categories. It is known in advance which of the third real data audio signal is associated with which of the fourth category of the audio signal. The additional criterion may be based on (or include) an age and / or a breed of a mammal (e.g., a dog) to which the audio signal belongs. That is, in order to train the third machine-learning-based classifier, only third real data audio signals corresponding to a specific age and / or breed and belonging to the first category but not the second category may be used.
[0043] In some embodiments, the input interface and / or the output interface is configured as a wireless interface. This enables the transmission of audio signals and / or results to and from the device. For example, the audio signal is sent to an application running on a smart phone or another portable device implementing the device for classifying audio signals.
[0044] According to a further aspect of the present invention, a method for classifying an audio signal is provided.
[0045] This method includes
[0046] ●receive input information of at least one audio signal;
[0047] ● classifying the audio signal into one of a first category and a second category of audio signals based on the input information of the audio signal by a first machine learning based classifier;
[0048] If the audio signal belongs to the first category of audio signals, then
[0049] classifying the audio signal into one of a plurality of third categories of audio signals based on the input information of the audio signal by a second machine learning based classifier; and
[0050] ●Output information about what category the audio signal belongs to.
[0051] In some embodiments, the method further includes: during a training phase, training the first machine learning-based classifier by real data audio signals including the first category and the second category of audio signals, so that the first machine learning-based classifier can determine whether the audio signal belongs to the first category or the second category of audio signals; and training the second machine learning-based classifier by real data audio signals including the first category but not the second category of audio signals, so that the second machine learning-based classifier can determine whether the audio signal belongs to one of the plurality of third categories, wherein it is known in advance which of the real data audio signals are related to which of the third categories of audio signals.
[0052] In some embodiments, the method includes classifying the audio signal into one of a plurality of fourth categories of audio signals based on the input information of the audio signal by a third machine learning based classifier if the audio signal belongs to one of the third categories of audio signals and satisfies an additional criterion.
[0053] The third machine-learning based classifier may be trained based on a third real data audio signal of the first category of audio signals but not the second category and satisfying the additional criterion, so that the third machine-learning based classifier is able to determine whether the audio signal belongs to one of the plurality of fourth categories. It is known in advance which of the third real data audio signals are associated with which of the fourth categories of audio signals. The additional criterion may be based on (or include) an age and / or a breed of a mammal (e.g., a dog) to which the audio signal belongs. That is, to train the third machine-learning based classifier, only third real data audio signals corresponding to a specific age and / or breed and belonging to the first category but not the second category may be used.
[0054] In some embodiments, the audio signal includes periodic heart sounds, wherein the first category of the audio signal represents a pathological heart murmur and the second category of the audio signal represents a healthy heart sound, wherein the plurality of third categories of the audio signal are associated with different pathological heart murmur severity levels.
[0055] According to a further aspect of the present invention, a computer program is provided. The computer program has a program code for executing the above method when the computer program is executed on a computer, a processor or a programmable hardware component.
[0056] As indicated above, one possible application of embodiments of the present invention is the diagnosis of a disease in an animal, such as a non-human mammal, in particular a dog. Thus, according to embodiments, the audio signal may be a recording of a heartbeat sequence of a dog or another animal or another non-human mammal and / or a recording of a heart murmur sequence of a dog or another animal or another non-human mammal. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Some examples of apparatus and / or methods will be described below, by way of example only, and with reference to the accompanying drawings, in which
[0058] Figure 1 A schematic block diagram of a device for classifying at least one audio signal according to the present invention is shown;
[0059] Figure 2 An example of an audio signal corresponding to a cycle of heart sounds is shown;
[0060] Figure 3 A flow chart showing a method for classifying at least one audio signal according to the present invention; and
[0061] Figure 4 A flow chart of a method for classifying at least one audio signal according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0062] Some examples are now described in more detail with reference to the accompanying drawings. However, other possible examples are not limited to the features of these embodiments described in detail. Other examples may include modifications of features and equivalents and alternatives of features. In addition, the terms used herein to describe certain examples should not limit further possible examples.
[0063] Throughout the description of the figures, the same or similar element symbols refer to the same or similar elements and / or features, which may be implemented in the same or modified form while providing the same or a similar function. The thickness of the lines, layers and / or regions in the figures may also be exaggerated for clarity.
[0064] When an "or" is used to combine two elements A and B, this is understood to disclose all possible combinations, i.e., only A, only B, and A and B, unless explicitly defined otherwise in individual cases. As an alternative wording for the same combination, "at least one of A and B" or "A and / or" B can be used. This applies equally to combinations of more than two elements.
[0065] If a singular form (such as "a / an" and "the") is used and the use of only a single element is not explicitly or implicitly defined as mandatory, further examples may also use multiple elements to implement the same function. If a function is described below as being implemented using multiple elements, further examples may use a single element or a single processing entity to implement the same function. It should be further understood that the terms "include / including" and / or "comprise / comprising" when used describe the presence of specified features, integers, steps, operations, procedures, elements, components and / or a group thereof, but do not exclude the presence or addition of one or more other features, integers, steps, operations, procedures, elements, components and / or a group thereof.
[0066] Figure 1 A block diagram of an apparatus 10 for classifying an audio signal is schematically shown. The apparatus 10 may be a programmable hardware device including memory and one or more processing units (such as a CPU and / or GPU). The apparatus 10 may be a mobile phone, a tablet computer, a personal computer, or the like. In other embodiments, the apparatus 10 may be implemented by a central server.
[0067] The device 10 includes an input interface 12 configured to receive at least one audio signal 20 or other characteristic information of the at least one audio signal. The audio signal 20 may be a digital representation of an acoustic signal. For example, the audio signal 20 may be received in a waveform audio file format (wav file). The audio signal 20 may, for example, be associated with a heartbeat of the heart of a mammal (e.g., a dog) and may therefore include a periodic or quasi-periodic audio pattern, such as a plurality of cycles of a periodic heart sound. A heart murmur is a unique heart sound produced when blood flows across a heart valve or blood vessel. In Figure 2 An example of an audio signal 20 corresponding to a cycle of heart sounds is shown in . For example, one cycle of a heart sound may be defined as the duration from signal (peak) S1 to the next subsequent signal (peak) S1. Subsequent cycles of a heart sound may typically differ slightly from each other in duration and / or amplitude of signal peaks S1, S2. Those skilled in the art having benefit of the present invention will appreciate that the audio signal 20 may also be associated with other acoustic signals, such as acoustic signals originating from, for example, a machine, without departing from the principles presented herein.
[0068] The input interface 12 of the device 10 can be a wired or wireless interface, such as a universal serial bus (USB), a WiFi interface, a Bluetooth interface, an infrared interface, or a cellular communication interface. In this way, the audio signal 20 (or other characteristic information thereof) can be conveniently transmitted from a data source to the device 10.
[0069] For example, apparatus 10 may optionally include one preprocessor 14 that may be configured to filter audio signal 20 in order to remove or reduce undesirable signal components, such as noise. Thus, preprocessor 14 may include one or more digital filters, such as low-pass, band-pass, or high-pass filters.
[0070] Additionally or alternatively, the optional preprocessor 14 may be configured to extract a plurality of features 22 that characterize the audio signal 20 from the audio signal 20. The preprocessor 14 may be configured to extract time domain features and / or frequency domain features from the audio signal 20, wherein the time domain and / or frequency domain features characterize or identify the audio signal 20. Examples of domain analysis methods include Fourier transforms, in particular fast Fourier transforms, power spectral density or wavelet decomposition transforms. Feature extraction starts with an initial set of measured data (audio signal 20) and establishes derived values (features) that are intended to be informative and non-redundant, thereby facilitating subsequent learning and generalization steps and, in some cases, leading to better human interpretation. Feature extraction is related to dimensionality reduction. Therefore, one dimension of the plurality of extracted features 22 is smaller than one dimension of the audio signal 20. The plurality of features 22 can also be regarded as feature vectors containing characteristic information of the audio signal.
[0071] The preprocessor 14 may be configured to automatically extract characteristic features of the audio signal 20, such as the heart sounds of a mammal. To this end, the audio signal 20 may be decomposed in a plurality of time intervals, each time interval comprising a heartbeat cycle. Thus, the feature extraction processor 14 may be configured to divide the audio signal 20 into a plurality of time intervals. For example, a determination of a time interval may be based on an algorithm that finds repetitions within the audio signal 20. For example, each time interval may include at least one cycle of a periodic audio pattern, such as a heartbeat cycle. The feature extraction processor 14 may be configured to extract features 22 of the audio signal that characterize the time interval from each time interval.
[0072] The interval length can be determined based on the duration of the audio signal 20 and the number of cycles of a periodic audio pattern. The calculation can be performed by a simple division. Of course, according to further embodiments, the interval length is determined differently, for example, by determining the duration of each cycle (for example, the time interval between a systole (diastole) and a subsequent systole (diastole)) and averaging these durations. According to further embodiments, the interval length can change over time, for example when the periodicity of the pattern changes. This can occur, for example, when the heart rate decreases in the current situation.
[0073] Possible time domain features that can be extracted from the audio signal 20 or an interval thereof are:
[0074] Mean of an audio pattern, median of an audio pattern, standard deviation of an audio pattern, variance of an audio pattern or relative to another pattern, skewness of an audio pattern, kurtosis of an audio pattern, mean absolute deviation of an audio pattern, 25th quantile of an audio pattern, 75th quantile of an audio pattern, entropy of an audio pattern, zero crossing rate of an audio pattern, quest factor, duration of a first peak, duration of another peak, duration from the end of S1 to the beginning of the next S1, duration from the end of S2 to the beginning of the next S1. In particular, when the audio signal 20 is divided into intervals, the duration feature is more meaningful.
[0075] Possible frequency domain features that can be extracted from the audio signal 20 or a section thereof are:
[0076] Mel-frequency cepstral coefficients, sound density, spectral flatness, spectral kurtosis, spectral skewness, spectral slope, spectral entropy, dominant frequency, bandwidth, spectral centroid, spectral flux and / or spectral roll-off.
[0077] It should be noted that the list of different feature types and the feature types are not limited to the list and feature types mentioned above. According to an embodiment, feature extraction can be performed mainly or completely automatically. In particular, compartmentalization can be performed automatically.
[0078] The apparatus 10 further includes a trained first machine-learning-based classifier 16 configured to map information of the audio signal 20 to one of a first class 24 and a second class 26 of the audio signal. In one embodiment, the information of the audio signal corresponds to the optionally extracted features 22, or in another embodiment, the audio signal 20 itself. In other words, the trained first machine-learning-based classifier 16 is configured to determine whether the audio signal 20 or its extracted features 22 is indicative of the first class 24 of the audio signal or the second class 26 of the audio signal.
[0079] In some embodiments, the first category 24 of audio signals represents an irregular audio signal, and the second category 26 of audio signals represents a regular audio signal. Here, an "irregular audio signal" may be understood as an abnormal audio signal. That is, an audio signal having a signal course that deviates from a normal or expected signal course. Conversely, a "regular audio signal" may represent an audio signal having a signal course that corresponds to a normal or expected signal course. In an example related to an audio signal 20 comprising heart sounds, the first category 24 of audio signals may indicate a pathological heart murmur, while the second category 26 of audio signals may indicate a healthy heart sound. Thus, the first category 24 of audio signals may indicate myxomatous mitral valve disease (MMVD), while the second category 26 of audio signals may indicate a healthy (non-human) mammal (e.g., a dog).
[0080] For example, the trained first machine learning based classifier 16 may be based on a neural network, a random forest algorithm or a boosting algorithm, in particular based on an XGboost algorithm.
[0081] The apparatus 10 further includes a trained second machine learning-based classifier 18 configured to map the extracted features 22 or the audio signal 20 belonging to the first category 24 of audio signals to one of the plurality of third categories 28A to 28C of audio signals. Thus, the trained second machine learning-based classifier 18 maps the extracted features 22 or the audio signal belonging to the first category 24 to category 28-A, category 28-B, or category 28-C. A skilled person having benefited from the present invention will appreciate that a different number (or amount) of third categories 28 is also possible. The trained second machine learning-based classifier 18 is only used for the audio signal 20 (or its characteristic features) belonging to the first category 24 of audio signals (e.g., pathological heart murmurs) and not for the audio signal 20 (or its characteristic features) belonging to the second category 26 of audio signals (e.g., healthy heart murmurs).
[0082] The plurality of third categories 28A-28C of audio signals may be associated with different levels of irregularity of the audio signal 20. That is, the plurality of third categories 28A-28C of audio signals may indicate how much the audio signal 20 deviates from a normal or expected audio signal (e.g., healthy heart sounds). For example, the plurality of third categories 28A-28C of audio signals may be associated with different levels of pathological heart murmurs. If the first category 24 of the audio signal indicates a pathological heart murmur or MMVD, for example, the plurality of third categories 28-A-28-C may be associated with different levels of severity of the pathological heart murmur (such as mild, moderate, or loud / throbbing). Thus, the output of the trained second machine learning-based classifier 18 can indicate whether the audio signal 20 contains a mild pathological heart murmur (category 28-A, mild deviation from healthy heart sounds), a moderate pathological heart murmur (category 28-B, moderate deviation from healthy heart sounds), or a loud / pathological heart murmur (category 28-C, strong deviation from healthy heart sounds).
[0083] Like the trained first machine learning based classifier 16, the trained second machine learning based classifier 18 can be based on, for example, a neural network, a random forest algorithm, or a boosting algorithm, in particular, an XGboost algorithm. Specifically, the trained first machine learning based classifier 16 and the second machine learning based classifier 18 can rely on the same type of machine learning algorithm or model. For example, the trained first machine learning based classifier 16 and the second machine learning based classifier 18 can both rely on the same boosting algorithm and also be programmed using the Python programming language.
[0084] However, the trained first machine learning-based classifier 16 and the second machine learning-based classifier 18 may be trained using different training signals (real data). During an initial training phase prior to the inference phase, the first machine learning-based classifier 16 may be trained by real data audio signals including both the first category 24 (e.g., pathological heart murmur) and the second category 26 (e.g., healthy heart sounds) of audio signals, so that the first machine learning-based classifier 16 is able to determine whether the audio signal 20 belongs to the first category 24 or the second category 26 of the audio signal. Usually for the real data, it is known in advance which of the real data audio signals (or their characteristic information) are related to which of the first category 24 and the second category 26 of the audio signal. Alternatively, the second machine learning-based classifier 18 may be trained by real data audio signals including only the first category 24 (e.g., pathological heart murmur) but not the second category (e.g., healthy heart sounds) of the audio signal, so that the second machine learning-based classifier 18 is able to determine whether the audio signal 20 belongs to one of the plurality of third categories 28A to 28C. Typically for real data, it is known in advance which of the real data audio signals (first category audio signals) of the second machine learning based classifier 18 are related to which of the third categories 28A to 28C of audio signals.
[0085] Due to different training data, the trained first machine-learning-based classifier 16 and the trained second machine-learning-based classifier 18 are configured differently. Specifically, the trained first machine-learning-based classifier 16 and the trained second machine-learning-based classifier 18 can be differently trained XGboost classifiers. The XGboost classifier is an example of a classifier based on a boosting algorithm.
[0086] The device 10 further includes an output interface (not shown) configured to output information about which of the categories 24, 26, 28A to 28C the audio signal 20 belongs to. For example, the output interface may include a display and / or a wired or wireless interface, such as a universal serial bus (USB), a WiFi interface, a Bluetooth interface, an infrared interface, or a cellular communication interface. In this way, output information can be conveniently transmitted from the device 10 to a data recipient.
[0087] Those skilled in the art who have benefited from the present invention will appreciate that the apparatus cannot be used to automatically classify only a single audio signal, but to automatically classify a plurality of audio signals. Depending on the implementation, the plurality of audio signals may be classified one by one or in parallel.
[0088] The apparatus 10 may be implemented by a single programmable hardware device or by a plurality of programmable hardware devices. Thus, the apparatus 10 is configured to perform a computer-implemented method 30 for classifying an audio signal. Figure 3 A flow chart of a computer-implemented method 30 is depicted in FIG.
[0089] The method 30 includes an act 32 of receiving at least one audio signal 20. The audio signal 20 may be associated with the heartbeat of a (non-human) mammal, such as a dog. Furthermore, the method 30 may include an act 34 of (automatically) extracting from the audio signal 20 a plurality of features 22 that characterize the audio signal 20. As previously mentioned, the features 22 may include time domain and / or frequency domain features that characterize the audio signal. The method 30 further includes an act 36 of classifying the audio signal 20 into one of a first category and a second category of audio signals by a first machine learning based classifier 16 based on the audio signal 20 or the extracted features 22. The first category 24 may be associated with pathological heart murmurs and the second category 26 may be associated with healthy heart sounds. A person skilled in the art having the benefit of the present invention will appreciate that the first category may also be associated with other abnormal sounds (e.g., abnormal machine sounds) and the second category 26 may also be associated with other normal sounds (e.g., normal machine sounds). Only in the case that the audio signal 20 belongs to the first category 24 of audio signals, the method 30 proceeds to an action 38 of classifying the audio signal 20 into one of a plurality of third categories 28 of audio signals by a second machine learning-based classifier 18 based on the extracted features 22 or the audio signal 20 itself. The third category 28 may be associated with different severity levels of pathological heart murmurs. Those skilled in the art who have benefited from the present invention will appreciate that the third category 28 may also be associated with different severity levels of other abnormal sounds (e.g., abnormal machine sounds). Still further, the method 30 includes an action 39 of outputting information about which category the audio signal 20 belongs to.
[0090] The method 30 may also include a further action of classifying the audio signal into one of a plurality of fourth categories of audio signals by a third machine learning based classifier. The classification may be based on the audio signal or features thereof and may be performed only if the audio signal belongs to one of the third categories of audio signals and satisfies an additional criterion. An embodiment of a machine learning based classifier with three sequences will be described below.
[0091] Now turn Figure 4 , shows a detailed flow chart of a computer-implemented method 40 for classifying an audio signal 20. For example, the method 40 may be run on a smartphone or a tablet computer in the form of an application.
[0092] Initially, a first machine learning based classifier 16 (e.g., a boosting algorithm), a second machine learning based classifier 18 (e.g., a boosting algorithm), and a third machine learning based classifier 42 (e.g., a boosting algorithm) are provided with previously trained first model parameters 41. This can be done in the form of one or more pickle files that can be used to save a machine learning model and serialize Python object structures. Once the first machine learning based classifier 16, 18, 42 is initialized, at least one audio file containing an audio signal 20 can be provided. For example, the audio signal 20 can indicate the heart sounds of a dog (or other mammal) recorded by a stethoscope. For example, the one or more audio signals 20 may appear as .wav files and may be provided via a wireless interface. In addition to the audio file(s), information about an age (integer) and a breed (string) of the dog from which the heart sounds originated may also be provided via the wireless interface.
[0093] The audio signal 20 extracted from the audio file can be provided to the preprocessor 14. The processor 14 may include a bandpass filter 14-1, a heartbeat cycle detection unit 14-2 and a feature extraction unit 14-3. For example, the bandpass filter 14-1 may have a bandpass from 50 to 500 Hz. In addition, the initial 0.5s and the last 0.5s of the audio signal 20 may be cut off to eliminate unwanted acoustic signals. The heartbeat cycle detection unit 14-2 is configured to perform the previously described binning to obtain one or more time intervals of the audio signal 20, wherein a time interval may correspond to at least one heartbeat cycle of the heart sound. The feature extraction unit 14-3 is configured to extract a plurality of characteristic features 22 from one or more intervals of the audio signal 20.
[0094] The characteristic feature 22 may then be provided as an input to the first machine learning based classifier 16, which is configured to map its input to a first model output indicating a first class 24 or a second class 26 of the audio signal. That is, depending on the characteristic feature 22, the first machine learning based classifier 16 will predict whether the audio signal 20 indicates a pathological heart murmur or a healthy heart sound. If the audio signal 20 indicates a healthy heart sound (class 26), this information may be output and the method 40 ends or returns. For example, a new audio signal may then be analyzed.
[0095] If, on the other hand, the audio signal 20 indicates a pathological heart murmur (class 24), the extracted characteristic features 22 of the audio signal 20 may be provided as input to a second machine learning based classifier 18, which is configured to map its input to a second model output indicating one of three classes 28-A, 28-B, and 28-C of the audio signal. The three classes 28-A, 28-B, and 28-C indicate mild, moderate, and loud / throbbery pathological heart murmurs. That is, depending on the input features 22, the second machine learning based classifier 18 will predict whether the audio signal 20 indicates mild, moderate, and loud / throbbery pathological heart murmurs. The respective information may be output and the method 40 may end or return.
[0096] If the age of the dog is less than or equal to a predefined age threshold (here: 2 years), the pathological heart murmurs may be classified as mild, moderate, and loud / thumping congenital heart murmurs. Congenital means that the dog was born with the condition. If the age of the dog is greater than the predefined age threshold (here: 2 years), this is classified as myxomatous mitral valve disease (MMVD) or dilated cardiomyopathy (DCM). The respective information may be output and the method 40 may end or return.
[0097] If the dog's breed indicates a small or medium dog, the pathological heart murmur is classified as mild, moderate, and loud / throbbing MMVD murmurs. If the dog's breed indicates a large dog, the pathological heart murmur is classified as mild, moderate, and loud / throbbing DCM murmurs. The respective information may be output and method 40 may end or return.
[0098] By training a third machine learning based classifier 42 (e.g., a boost algorithm), an audio signal 20 classified as one of mild, moderate, and loud / throbbing MMVD murmurs is classified as one of two MMVD stages of heart disease. Here, the two stages are B1 and B2. Stage B1 describes asymptomatic dogs with no radiological or echocardiographic evidence of cardiac remodeling in response to their MMVD, and dogs with remodeling changes that are present but not severe enough to meet current clinical trial criteria for determining the need for initiation of treatment. Stage B2 refers to asymptomatic dogs with more advanced mitral regurgitation, with mitral regurgitation hemodynamics that are severe and long-standing, sufficient to cause radiographic and echocardiographic findings of left atrial and ventricular enlargement, which meet clinical trial criteria for identifying dogs that should clearly benefit from initiation of drug treatment to delay the onset of heart failure.
[0099] If the audio signal 20 indicates a pathological heart murmur (category 24), one of mild, moderate, and loud / throbbing pathological heart murmurs (one of three categories 28-A, 28-B, and 28-C), the dog is over 2 years old, and its breed is small / medium, then the extracted characteristic features 22 of the audio signal 20 can be provided as input to a third machine learning-based classifier 42, which is configured to map its input to a third model output indicating one of two categories 44-A, 44-B. The two categories 44-A, 44-B indicate mild, moderate, and loud / throbbing MMVD murmurs in phase B1 or B2. That is, depending on the input features 22, the second machine learning-based classifier 18 will predict whether the audio signal 20 indicates mild, moderate, and loud / throbbing MMVD murmurs in phase B1 or B2. The respective information can be output and the method 40 can end or return.
[0100] It should be noted that the third machine learning-based classifier 42 can be trained based on third real data audio signals that are of the first category 24 (pathological heart murmur) but not of the second category and that satisfy the additional criteria of dog age>2 and dog breed=small / medium. That is, to train the third machine learning-based classifier 42, only third real data audio signals that correspond to a specific age and / or breed and belong to the first category 24 but not of the second category can be used.
[0101] The embodiment of the present invention proposes a concept that intensity can be detected in addition to sound detection. For this purpose, a multi-part / cascade algorithm is proposed: a trained model (here a feature matrix) can be loaded into the algorithm in the form of a pickle file. Each new sample (audio signal) to be tested goes through the same steps as the trained model. New samples / sounds can be filtered out from noise (environmental noise, ...). For this purpose, different methods of time-frequency analysis can be run (fast Fourier transform, power spectral density, wavelet decomposition). Subsequently, heart sounds can be detected by using different methods so that several intervals each having a heartbeat cycle can be obtained. Various features 22 can be calculated for individual heartbeat cycles (intervals) in both the time domain and the frequency domain. The same steps were previously performed on a pickle file with multiple data. Finally, a new test data set can be given to the first classifier 16. This uses all the data and checks whether this is a pathological heart murmur or a healthy heart sound. Predicted health data is not used for the second classifier 18. Therefore, only pathological data sets (i.e., only a subset of the data set used at the beginning) will enter the second classifier 18. This classifier 18 then checks the loudness of the heart murmur and returns whether it is a mild, moderate or loud murmur. Using this information, a recommendation can be issued to a physician. Dogs under 2 years old may have a congenital murmur. Dogs over 2 years old with a small / medium breed may have MMVD disease. Large dogs usually have DCM (dilated cardiomyopathy). Boosting algorithms (such as XGBoost, AdaBoost, etc.) are particularly well suited for this.
[0102] Aspects and features described with respect to a particular one of the preceding examples may also be combined with one or more of the further examples to replace the same or similar features of the further examples or to otherwise introduce features into the further examples.
[0103] Examples may further be or relate to a (computer) program comprising a program code to perform one or more of the above methods when the program is executed on a computer, processor or other programmable hardware component. Therefore, the steps, operations or procedures of different ones of the methods described above may also be performed by a programmed computer, processor or other programmable hardware component. Examples may also cover program storage devices, such as digital data storage media, which are readable by machines, processors or computers and encode and / or contain machine-executable, processor-executable or computer-executable programs and instructions. For example, the program storage device may include or be a bit storage device, a magnetic storage medium (such as a disk and tape), a hard drive or an optically readable digital data storage medium. Other examples may also include a computer, a processor, a control unit, a (field) programmable logic array ((F)PLA), a (field) programmable gate array ((F)PGA), a graphics processor unit (GPU), an application-specific integrated circuit (ASIC), an integrated circuit (IC) or a system-on-chip (SoC) system programmed to perform the steps of the methods described above.
[0104] It should be further understood that the disclosure of several steps, procedures, operations or functions disclosed in the description or in the scope of the invention patent application should not be interpreted as implying that such operations must depend on the described order, unless explicitly stated in individual cases or required for technical reasons. Therefore, the previous description does not limit the execution of several steps or functions to a specific order. In addition, in further examples, a single step, function, procedure or operation may include and / or be decomposed into several sub-steps, sub-functions, sub-routines or sub-operations.
[0105] If aspects are described with respect to a device or system, these aspects should also be understood as a description of a corresponding method. For example, a block, device, or functional aspect of a device or system may correspond to a feature (such as a method step) of a corresponding method. Therefore, aspects described with respect to a method should also be understood as a description of a property or a functional feature of a corresponding block, a corresponding element, a corresponding device, or a corresponding system.
[0106] The following claims are hereby incorporated into the detailed description, where each claim may be independently presented as a separate example. It should also be noted that although in the claims, a dependent item is referred to in a specific combination with one or more other claims, other examples may also include a combination of the dependent item with the subject matter of any other dependent or independent item. Such combinations are hereby expressly proposed unless they are stated in individual cases where a specific combination is not intended. In addition, features of a claim shall also be included with respect to any other independent item, even if the claim is not directly defined as dependent on the other independent item.
Claims
1. A device (10) for classifying at least one audio signal (20), the device (10) comprising an input interface (12) configured to receive input information (22) of the audio signal (20); A first machine learning based classifier (16) is trained and configured to map the input information (22) of the audio signal to one of the first and second classes (24; 26) of audio signals; A second machine learning based classifier (18) is trained, which is configured to If the audio signal (20) belongs to the first category (24) of audio signals, then Mapping the input information (22) of the audio signal belonging to the first category (24) of audio signals to one of a plurality of third categories (28) of audio signals; and An output interface is configured to output information about which category the audio signal (20) belongs to.
2. The device (10) as claimed in claim 1, wherein: The audio signal (20) includes a plurality of cycles of heart sounds.
3. The device (10) as claimed in claim 2, wherein: The first category (24) of audio signals represents pathological heart murmurs, and the second category (26) of audio signals represents healthy heart sounds.
4. The device (10) according to claim 2 or 3, wherein: The plurality of third categories (28) of audio signals are associated with different levels of pathological heart murmurs.
5. The device (10) as claimed in any one of the preceding claims, comprising: A preprocessor (14) is configured to extract a plurality of features (22) characterizing the audio signal from the audio signal (20) as the input information (22) of the audio signal (20).
6. The device (10) as claimed in claim 4, wherein: The preprocessor (14) is configured to extract time domain features and / or frequency domain features that characterize the audio signal (20).
7. The device (10) as claimed in any one of the preceding claims, wherein: The trained first machine learning based classifier (16) is configured to implement a trained first boosting algorithm, and / or wherein the trained second machine learning based classifier (18) is configured to implement a trained second boosting algorithm.
8. The device (10) as claimed in any one of the preceding claims, wherein: The trained first and second machine learning based classifiers (16; 18) are of the same type and are distinguished by different respective training signals.
9. The device (10) as claimed in claim 8, wherein: The first machine learning-based classifier (16) is trained based on a first real data audio signal including the first category (24) and the second category (26) of audio signals, so that the first machine learning-based classifier (16) can determine whether the audio signal (20) belongs to the first category or the second category of audio signals, wherein it is known in advance which of the first real data audio signals is related to which of the first and second categories of audio signals, and wherein the second machine learning-based classifier (18) is trained based on a second real data audio signal including the first category (24) of audio signals but not the second category, so that the second machine learning-based classifier (18) can determine whether the audio signal (20) belongs to one of the plurality of third categories (28), wherein it is known in advance which of the second real data audio signals is related to which of the third categories of audio signals.
10. The device (10) as claimed in any one of the preceding claims, further comprising: A trained third machine learning based classifier (42) is configured to map the input data (22) of the audio signal belonging to any one of the plurality of third classes (28) of audio signals and satisfying an additional criterion to one of the plurality of fourth classes of audio signals.
11. The device (10) according to claim 10, wherein: The trained first, second and third second machine learning based classifiers (16; 18; 42) are of the same type and are distinguished by different respective training signals.
12. The device (10) according to claim 10 or 11, wherein: The third machine learning-based classifier (42) is trained based on a third real data audio signal of the first category (24) of the audio signal but not of the second category and satisfying the additional criterion, so that the third machine learning-based classifier (42) is able to determine whether the audio signal (20) belongs to one of the plurality of fourth categories (28), wherein it is known in advance which of the third real data audio signals is related to which of the fourth categories of the audio signal.
13. The device (10) according to any one of claims 10 to 12, wherein: The additional criterion is based on the age and / or species of the mammal to which the audio signal (20) belongs.
14. A method for classifying at least one audio signal, the method comprising Receiving input information (22) of the audio signal (20); classifying the audio signal (20) into one of first and second categories (24; 26) of audio signals based on the input information (22) by a first machine learning based classifier (16); If the audio signal (20) belongs to the first category (24) of audio signals, then classifying the audio signal (20) into one of a plurality of third categories (28) of audio signals based on the input information (22) of the audio signal by a second machine learning based classifier (18); and Information about the category to which the audio signal (20) belongs is output.
15. The method of claim 14, further comprising: During the training phase, The first machine learning-based classifier (16) is trained by using real data audio signals including the first category (24) and the second category (26) of audio signals, so that the first machine learning-based classifier (16) can determine whether the audio signal (20) belongs to the first category or the second category of audio signals; and The second machine learning-based classifier (18) is trained by using a real data audio signal that includes the first category (24) of audio signals but not the second category, so that the second machine learning-based classifier (18) is able to determine whether the audio signal belongs to one of the plurality of third categories (28), wherein it is known in advance which of the real data audio signals is related to which of the third categories of audio signals.