Method for classifying digital audio data that follow each other in time
By converting digital audio data in road traffic into frequency expressions and using classifiers for classification, the problem of difficult to identify and classify acoustic special signals in the prior art is solved, and the recognition effect of high sensitivity and low error rate is achieved.
Patent Information
- Application Number
- CN202010298706.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-04-17
- Filing Date
- 2020-04-16
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2040-04-16
AI Technical Summary
The prior art is difficult to effectively identify and classify acoustic special signals in road traffic, especially when the error rate is close to zero, resulting in unnecessary or erroneous reactions in vehicles, such as red lights and traffic delays.
By converting digital audio data that follows each other in time into frequency expressions, and calculating multiple frequency expressions using short-time Fourier transform or wavelet transform, the acoustic signals representing dangerous situations are then classified through a classifier.
High sensitivity recognition of acoustic special signals is achieved, error rates are reduced, unnecessary vehicle reactions are avoided, and special signals are distinguished from different countries.
Smart Images

Figure CN111833904B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a method for classifying digital audio data which describe acoustic signals which characterize dangerous situations such as occur, for example, in road traffic. Background Art
[0002] Until now, there have been no systems with special signal recognition in road traffic, since practical applications require systems with high sensitivity, which must ensure that acoustic signals in road traffic are classified with respect to the presence of special signals with a negligible number of false alarms. Since the use of such systems in road traffic can only be justified if the error rate is very close to zero, in order to avoid unwanted or possibly erroneous reactions of vehicles, such as running a red light, and the associated delays in road traffic. Such systems should also be able to distinguish between the different special signals used worldwide. Summary of the invention
[0003] The invention discloses a method for classifying temporally successive digital audio data, which describe acoustic signals characterizing a dangerous situation, as well as a computer program, a computer-readable storage medium and a decision-making system according to the features of the independent claims. Advantageous embodiments are the subject matter of the dependent claims and the following description.
[0004] Not only for driver assistance systems, but also in the field of at least partially automated driving, it is important to detect emergency vehicles with special signals and traffic police with acoustic signal transmitters according to legal regulations in various countries. Other acoustic signals that represent dangerous situations, such as emergency calls or warning signals from other vehicles, should also be detected so that corresponding actions can be initiated, if necessary automatically, or the situation can be indicated to the vehicle driver.
[0005] Furthermore, at least partially automated vehicles with acoustic special signal recognition offer the advantage that by identifying such audio signals, an early situation assessment is possible for the vehicle driver and for the partially automated system, even when a special vehicle or traffic police officer is not yet directly visible, so that they can react accordingly.
[0006] For at least partially automated vehicles, when approaching a difficult-to-see intersection or in an emergency, a suitable driving route can be selected early with the help of the correspondingly classified audio signal information in order to give way to emergency vehicles. Alternatively, if no recognition system is used within the scope of at least partially automated driving, a corresponding notification of the driver can take place when a special signal is recognized. This is particularly advantageous for hearing-impaired people in driver assistance systems, but can also be advantageous for all vehicle drivers due to the higher sensitivity due to the sound transducer being arranged as far as possible outside the interior.
[0007] The privileges of the emergency vehicle with respect to other vehicles involved depend, for example, on the formation of an emergency lane, such as giving way or prohibiting entry to an intersection. Therefore, a suitable driving route must be identified. However, in different countries, different signals are used as special signals. For example, the following signal in Germany ("Martinshorn, siren") or the "Wail", "Yelp" or "Rumbler" in the United States.
[0008] The invention is based on the knowledge that, when identifying, in particular, special signals in acoustic situations of road traffic and other acoustic signals characterizing dangerous situations, an analysis of overtones belonging to a fundamental tone generally improves the identification of the signal to be determined and, in certain acoustic situations, even more clearly makes the overtone stand out from the background noise as the associated fundamental tone.
[0009] Acoustic signals that characterize dangerous situations are of great importance in particular in road traffic, since each traffic participant must react to such signals depending on the situation in which he or she finds himself or herself in road traffic. As important examples of such signals, special signals of emergency vehicles are mentioned, which are typically implemented as special sirens and generate different tone sequences, which are also country-specific. Examples of special signals are alarm bells: the tone sequence "Ta-Tü-Ta-Ta" with two basic tones within 3.0+ / -0.5s (between approximately 360 Hz and 630 Hz); as well as Wail, Yelp and Rumbler, which, however, must be distinguished from other tone sequences (English: change of constant notes, pitch change), distorted siren signals (the distorted siren signals, for example, come from toys or smartphones), and steady-state sirens (English: Civil defense siren, defense alarm). Other examples are car horns, which can come from different vehicle types, such as passenger vehicles, freight vehicles or trains or trams. Also important for road traffic users are acoustic signals from traffic police, such as whistles, acoustic bell signals at railway crossings, acoustic warning signals from vehicles driving back, and acoustic signals from automatic alarm systems. In addition, emergency calls or other calls, such as "Stop" or "Fire", must be taken into account in different languages and must be distinguishable from normal conversation or music, if necessary.
[0010] Such acoustic signals characterizing a dangerous situation can be converted into electrical signals by means of one or more sound transducers, such as microphones, wherein the sound transducers can be coupled acoustically as directly as possible to the sound environment to be monitored, such as a road traffic situation. For example, an arrangement of the sound transducer outside the vehicle cabin has the advantage over an arrangement in the interior area of the vehicle that the acoustic signal is not attenuated by the cabin boundary and is thus coupled directly to the sound environment.
[0011] From such electrical signals of the sound transducer, temporally subsequent digital audio data can be generated, for example by means of an electronic analog-to-digital converter circuit, which then contain the corresponding signal characterizing the danger situation in a digitally encoded manner.
[0012] The digital-to-analog conversion of the electrical signal can be performed in this case such that the acoustic signal characterizing the hazardous situation includes a significant frequency range of, for example, 250 Hz to 8 kHz, so that the electrical signal is detected with the aid of a sampling rate or sampling rate that is twice as high as the highest frequency to be detected and converted into digital audio data. A higher sampling rate can increase the accuracy of the conversion.
[0013] The above steps for providing digital audio data describing an acoustic signal indicative of a hazardous situation are primarily intended to illustrate, introduce, and define terminology.
[0014] The method according to the invention for classifying temporally successive digital audio data describing acoustic signals characteristic of a dangerous situation calculates in one step a plurality of frequency representations for progressively progressive time intervals of the temporally successive audio data.
[0015] A frequency representation is calculated for time intervals of digital audio data that follow one another in time, i.e., are staggered with respect to subsequent time intervals, wherein the calculation of the frequency representation can be performed with the aid of a plurality of alternative methods. An exemplary method is the so-called short-time Fourier transform (STFT), a further possible method is the so-called wavelet transform. The method is further explained in detail below. The calculation results in a spectrum, i.e., the amplitude of the frequency components of the acoustic signal described by the digital audio data with respect to the frequencies.
[0016] The progressive time intervals can overlap in time according to the method of the invention. A high degree of overlap of the time intervals results in a representation of the spectrum, ie the frequency representation of the digital audio data, with respect to time, with a high temporal resolution.
[0017] As an example, the calculation of the frequency representation can be calculated with the aid of a short-time Fourier transform with 2 to the 11th power 11 = 2048 digital audio samples, but it is understood by those skilled in the art that a number of other values for the audio samples are possible here. If the digital audio signal is generated with the aid of a sampling rate of 10 kHz, the data includes a time interval of 0.2048 seconds, and when the progression of the time interval takes place with a time interval of 0.1 seconds, a high detection accuracy is obtained with a small delay. As a result, the time intervals overlap with each other by approximately 50%. Depending on the signal to be detected, the required accuracy and the delay time for classification, the time intervals can also be adjusted, for example, in a range of 0.05 seconds to 0.2 seconds, for example, and the overlap can also be selected to be larger or smaller. With a time interval with a step size of 0.05 seconds, an improved classification performance is achieved for rapidly changing signals, such as "Yelp", by means of an increased time resolution.
[0018] The plurality of frequency representations may include, for example, 28 frequency representations, so that in each 0.1 second time interval a time region of approximately 3 seconds is included, which when detecting particularly important signals, for example a follow-up signal (Folgensignal) with a repetition frequency of 3 seconds follows a follow-up special signal of an emergency vehicle, so that the characteristic time variation of the signal can be analyzed. The values mentioned by way of example can be easily adapted to other classification tasks.
[0019] In a further step of the method, the individual frequency representations are divided into octaves, i.e. frequency ranges, the final frequency of which is twice the initial frequency. In each octave, a certain number of frequency segments of each individual frequency representation is formed, wherein the frequency segments comprise subsets of the individual frequency representations. The certain number of frequency segments may be, for example, 12, but any other number of divisions which are particularly beneficial for subsequent analysis of acoustic signals to be identified, for example, by classification, may be selected. Here, the division may also be based on a multi-resolution filter bank or on other merging strategies, wherein, for example, wider segments are used in the case of higher frequencies.
[0020] The corresponding frequency segments of each individual frequency representation of different octaves are added in another step according to the invention. The frequency segments are present in all octaves of the frequency representation, are sorted in the associated octaves according to the increased frequency and can be distributed in the same proportion to the corresponding octaves. Thus, the frequency segments of one octave correspond to the frequency segments of another octave, which are sorted in the same way in their octaves.
[0021] This allows the harmonic harmonics of the fundamental tone to be added to the signal component of the fundamental tone in the case of a coordinated signal formation, thereby having a larger value that can stand out more easily from the basic noise, such as traffic. If the fundamental tone is rarely present in the audio data, the signal to be detected is recognized due to the frequently present harmonics.
[0022] In a further step of the method, the frequency components are calculated by forming a mean value for the individual added frequency segments in each individual frequency representation.
[0023] Thus, by gradually reducing the amount of data to be processed, in addition to a simplified further processing of the data, it is achieved that small fluctuations in the frequency of the signal to be identified, which is described by the digital audio data, do not impair the quality of the data for identification and classification, or that noise components are suppressed for the analysis. The signal to be identified is reduced by the method by, for example, differentiation or variability of different generator types for special signals, which are based, for example, on pneumatic or electrical operating principles, as well as multiple superpositions of the signals to be identified, which significantly reduces the complexity of the classification task.
[0024] In a further step of the method, a classification vector is generated by means of a classifier and the number of frequency components of the plurality of frequency representations. For this purpose, the classifier is designed to classify a signal characterizing a dangerous situation and described by means of temporally successive digital audio data by means of the associated number of frequency components of the plurality of frequency representations and to associate the classified classification vector with a corresponding value.
[0025] With the described method according to the invention, a high sensitivity, ie, a radius of effectiveness, is thus achieved for the recognition of acoustic special signals and for other acoustic signals characteristic of a dangerous situation, which leads to a significantly earlier recognition of these signals.
[0026] Since the method allows the signal to be detected to stand out from the acoustic environment, a low error rate results, which is caused by the specific signal processing, in particular by superposition of frequency segments, together with a specially defined classifier. Since, for example, a siren type that is area-invalid does not trigger an unnecessary and erroneous reaction of the system when a special signal is detected, the disruption of the traffic flow can be reduced to a minimum. The method thus detects, for example, acoustic special signals and distinguishes between different special signal types.
[0027] According to one embodiment of the method, the frequency segments of the number of frequency segments for each octave of each individual frequency expression are arranged within the octave so that the frequency segments at least partially overlap each other. This superposition can also be carried out with the help of frequency segments that have been convolved beforehand with the help of a distribution function so that the contribution of the superposition area with other segments is not so strongly attenuated, even when the frequency superposition area is large.
[0028] The inclusion of a larger frequency range achieved thereby can result in a more robust classification and a better signal-to-noise ratio.
[0029] According to another design of the method, digital audio data that follow one another in time are analyzed with respect to the presence of fundamental tones and overtones, and before multiple frequency expressions are calculated for progressively progressive time intervals of the audio data that follow one another in time, frequencies in a frequency band around the fundamental tones and overtones in the audio data are attenuated.
[0030] By means of the preprocessing of the audio data, the audio data has an improved signal-to-noise ratio, in particular for special signals, such as special signals of emergency vehicles, and can be classified more robustly. For example, a frequency band of a certain width around the fundamental frequency and the first three harmonics of the fundamental frequency can be filtered out of the audio signal. In other words, the frequency components outside the frequency band are attenuated in order to filter out the frequency band.
[0031] Identifying fundamental frequencies and / or overtones in an audio signal can be performed with the aid of a series of analysis methods. For example, some methods are mentioned below, which illustrate corresponding features of the design of the method.
[0032] Pitch detection or fundamental frequency analysis (English: Pitch Detection) examines the time signal for the presence of prominent signal components by means of automatic correlation calculations of the audio signal. The filtering or reduction of frequency components outside the identified frequency band can then be performed by means of a bandpass filter.
[0033] Cepstrum analysis is a method based on the Fourier transform. The calculation is performed by taking the complex logarithm of the Fourier transform and a subsequent inverse Fourier transform. With the help of a bandpass filter, frequency components outside the identified frequency band can be filtered out.
[0034] Spectral flatness can be used, in particular in special signals to be classified, by their high spectral energy density at only individual discrete frequencies. In this case, only those frequency components that show peaks in the spectrum are retained, while all others are filtered out. "Spectral flatness" is a measure for "peaks / tones" and is calculated as follows: Spectral flatness = (geometric mean of the power spectrum) / (arithmetic mean of the power spectrum). In this case, no additional steps for filtering are necessary.
[0035] The "Adaptive Spectral Subtraction" filter is an adaptive filter, which can adjust its filter characteristics. In the case of a particularly noisy background, the filter can adapt to the current data and suppress the background more strongly. In this case, the background in the spectrum is calculated in regular time and frequency intervals, such as a flat background below a tone / peak-like frequency component, such as the "Ta-Tü-Ta-Ta" of a police bell, and subtracted from the overall spectrum.
[0036] Another possibility for preprocessing is to use an autoencoder neural network which learns a model in order to generate a compressed or de-noised representation of the input data by correspondingly extracting important features, in our case the audio signal from the general background.
[0037] An "autoencoder" is understood to be an artificial neural network KNN, which is able to learn certain patterns contained in the input data. An autoencoder is used to generate a compressed or de-noised representation of the input data by extracting important features such as certain classes, in our case the audio signal from the general background. An autoencoder uses three or more layers:
[0038] Input layer, such as a 2D image.
[0039] • Multiple significantly smaller layers that form the encoding to reduce data.
[0040] An output layer, the dimension of which corresponds to the dimension of the input layer, i.e., each output parameter in the output layer has the same meaning as the corresponding parameter in the input layer.
[0041] According to a further embodiment of the method, it is provided that a plurality of frequency representations are calculated for progressively successive time intervals of the audio data which follow one another in time by means of a short-time Fourier transformation or a wavelet transformation.
[0042] The time-limited or short-time Fourier transform (STFT=Short-Time Fourier Transform) is a procedure that provides a Fourier transform for non-stationary data. A Hanning window (Hanning window) is applied to the observed digital audio data, which reduces the beginning and end of the audio data to the value zero in order to reduce leakage effects and increase the time resolution. Each single fast Fourier transform (FFT) is associated with a time that corresponds to the middle of the window. The short-time Fourier transform with a window function has a fixed frequency-time resolution.
[0043] In wavelet analysis similar to STFT, a time-limited "wave packet" function is used instead of the infinitely extended sine / cosine function. The term wavelet transform WT represents a family of linear time-frequency transforms. Here, WT consists of the so-called wavelet analysis, i.e. the transition from time representation to spectrum representation, and wavelet synthesis, i.e. the inverse transformation of the wavelet transform into a time period. The wavelet transform has a high frequency resolution at low frequencies, but a low time localization. At high frequencies, the wavelet transform has a low frequency resolution, but a good time localization.
[0044] In particular, the calculation of the frequency representation by means of a short-time Fourier transform has the advantage that a particularly rapid calculation of the Fourier transform is carried out.
[0045] In one embodiment of the method, it is provided that the frequency components of each individual frequency expression are normalized before the classification vector is generated.
[0046] In a further embodiment of the method, it is provided that the frequency components of each individual frequency representation are normalized to the value one.
[0047] The advantage of normalization is that even less intense acoustic signals, which initially only stand out slightly from the noise of the digital audio data, stand out from the background by the normalization for the classifier.
[0048] In a variant of the method, it is proposed that the frequency components of each individual frequency representation are calculated by means of a histogram equalization method ("Histogram Equalization" in English). In this case, such scale values that occur less frequently, such as grayscale or color values of an image, are enhanced, while such scale values that occur particularly frequently are weakened. Compared to simple normalization to the maximum value, structures or contrasts in the data can be highlighted and enhanced in a targeted manner by means of the histogram equalization method.
[0049] According to another embodiment of the method, the classifier has an artificial neural feedforward network, which is designed and trained to classify the number of frequency components of a signal representing a dangerous situation described by means of temporally successive digital audio data expressed by means of multiple frequencies by generating the value of a classification vector.
[0050] The classification of special signals can be performed with the aid of an artificial neural network (KNN, in English: Artificial Neural Network ANN). KNN consists of a network of artificial neurons, which is a biological model, ie the networking of neurons in the nervous system / brain is modeled accordingly.
[0051] Neural networks provide a framework for multiple different algorithms to work together and process complex data inputs for machine learning. Such neural networks learn to perform tasks based on examples, and are typically not programmed with task-specific rules.
[0052] This neural network is based on a collection of connected units or nodes, which are called artificial neurons. Each connection can transfer a signal from one artificial neuron to another artificial neuron. The artificial neuron that receives the signal can process the signal and then activate other artificial neurons connected to it. In the conventional implementation of the neural network, the signal at the connection of the artificial neuron is a real number, and the output of the artificial neuron is calculated by a nonlinear function of the sum of its inputs. The connection of the artificial neuron typically has a weight, which is adjusted as the learning progresses. The weight increases or decreases the intensity of the signal at the connection. The artificial neuron may have a threshold value so that the signal is output only when the total signal exceeds the threshold value. Typically, multiple artificial neurons are integrated in layers. Different layers may perform different types of transformations for their inputs. The signal may move from the first layer, the input layer, to the last layer, the output layer after passing through the layers multiple times.
[0053] The structure of the artificial neural feedforward network may be a structure that is configured such that it receives a single data pattern according to an image at its input stage and provides an output classification vector containing a recognition probability for each class of interest.
[0054] According to a further embodiment of the method, it is proposed that the classifier has a multilayer perceptron (MLP) which is designed and trained to classify the number of frequency components of a signal characterizing a danger situation described by means of temporally successive digital audio data expressed by means of a plurality of frequencies by generating the value of a classification vector.
[0055] This network belongs to the family of feedforward artificial neural networks. In principle, an MLP consists of at least 3 layers of neurons: an input layer, an intermediate layer (hidden layer) and an output layer. This means that all neurons of the network are divided into layers, where the neurons of one layer are always connected to all neurons of the next layer. There are no connections to the previous layer and no connections that skip layers. In addition to the input layer, the different layers consist of neurons that are subjected to nonlinear activation functions and are connected to the neurons of the next layer.
[0056] According to another embodiment of the method, the classifier has an artificial neural feedback network, which is designed and trained to classify the number of frequency components of a signal representing a dangerous situation described by digital audio data that follows each other in time by means of multiple frequency expressions by generating the value of a classification vector. A feedback neural network (RNN in English) is a neural network that, compared to a feedforward network, also has connections between neurons of a layer and neurons of the same layer or a previous layer. The structure is particularly suitable for finding time-coded information in data.
[0057] According to another design of the method, the classifier has an artificial neural convolutional network, which is designed and trained to classify the number of frequency components of a signal representing a dangerous situation described by digital audio data that follow each other in time by means of multiple frequency expressions by generating the value of a classification vector.
[0058] In addition to the embodiment of the feedforward neural network above, the structure of the artificial neural convolutional network (ConvolutionalNeural Network) is composed of one or more convolutional layers followed by a pooling layer. The order of the layers can be used with or without a normalization layer (e.g., batch normalization), a padding layer, a dropout layer, and an activation function, such as a linear rectifier unit ReLu, a sigmoid function, a tanh function, or a softmax function.
[0059] In principle, the units can be repeated as often as desired, and if repeated sufficiently, we refer to a deep convolutional neural network. When repeated blocks consisting of convolutional and pooling layers are combined, the CNN ends with one (or more) fully connected layers, a structure similar to an MLP.
[0060] The structure of such a neural convolutional network is typically constructed from two parts. The first part is a sequence of layers that scan down the input grid at a lower resolution to obtain the desired information and store redundant information. The second part is a sequence of layers that scan the output of the first part back into the fully connected layer and produce the desired output resolution, such as a classification vector with the same length as a number of signals characterizing the dangerous situation to be classified.
[0061] According to another design of the method, the classifier is designed to classify the signal representing the dangerous situation by gradually comparing the number of frequency components of multiple frequency expressions with the number of templates, and to associate the classification vector with different values according to the comparison result.
[0062] The advantage of this simple approach is that, in contrast to KNN, the mode of operation of the algorithm can be more clearly understood.
[0063] According to another embodiment of the method, the classifier has a "support vector machine" SVM, which is designed and trained to classify the number of frequency components expressed by multiple frequencies of a signal characterizing a dangerous situation described by digital audio data that follows one another in time by generating the value of a classification vector.
[0064] The algorithm uses training data, in which it is known which class the training data belongs to and can be used as a classifier and as a regressor. The SVM divides the collection of data points / objects into classes in n-dimensional space so that around the class boundaries, the largest possible n-dimensional "area" remains free. So-called hyperplanes are therefore sought, which separate the data sets of different classes from each other as well as possible.
[0065] According to another design of the method, the classifier is a k-nearest-neighbor k-NN classifier. This is a non-parametric method for estimating a probability density function. Class association is performed only under the condition of considering k nearest neighbors. In the simplest case, classification is performed by a simple multiple decision, in which the k nearest objects participate. Object x is assigned to the following category, which has the largest number of objects among these k neighbors. In order to determine the k nearest neighbors, multiple spacing metrics (such as Euclidean spacing, etc.) can be considered. For this, a k-NN classifier is trained based on data of known categories.
[0066] According to another design of the method, the classifier has a pre-classifier and a main classifier, and the pre-classifier is designed to identify the signal representing the dangerous situation described by the digital audio data that follows each other in time by the number of frequency components expressed by multiple frequencies. The main classifier is designed to classify the signal representing the dangerous situation described by the digital audio data that follows each other in time by the number of frequency components expressed by multiple frequencies through the value of the classification vector if the pre-classifier has identified the signal representing the dangerous situation in the digital audio data that follows each other in time.
[0067] Since the pre-classifier has to solve a less complex task, this results in the advantage that it can detect more quickly and with fewer resources whether a signal representing a dangerous situation is present at all, before the more complex task of exact classification is connected.
[0068] According to another design of the method, the classifier has a plurality of sub-classifiers, which are respectively trained only for one signal to be classified and complete the classification task in parallel. In this way, the sub-classifiers can be trained individually with high specificity and are more robust to misclassification. For this, N classifiers are trained individually, where N is the number of alternative categories. The individual "binary" classifications can be combined into a total classifier in different ways. The final evaluation of the classification can then be performed by the category with the highest probability in the individual classifiers.
[0069] According to a further embodiment of the method, provision is made for an actuation signal for actuating the at least partially automated vehicle and / or a warning signal for warning a vehicle occupant to be emitted as a function of at least one of the values of the classification vector.
[0070] Based on the control signal, in particular the vehicle can be guided longitudinally or transversely.
[0071] Such an actuation signal can thus be supplied, for example, to a control unit or an actuator, which can then respectively initiate an operation, such as a steering operation, an acceleration or a braking operation.
[0072] Based on the warning signal, for example, the display unit can be actuated in such a way that the vehicle occupants receive information about upcoming events, such as the approach of an emergency vehicle, so that a controlled operation can be carried out by the driver.
[0073] The above-described design of the method provides the following advantages, which improve safety in road traffic: On the one hand, the emergency vehicle can move forward faster, and on the other hand, traffic accidents caused by the emergency vehicle can be prevented.
[0074] An at least partially automated vehicle may also be understood to be a robot, such as a logistics and / or industrial robot. Mobile garden tools, such as lawn mowers, etc., which are operated at least partially in an automated manner, also fall within this definition.
[0075] The at least partially automated vehicle may also be another mobile robot, for example a mobile robot that moves forward by flying, swimming, diving or walking. The mobile robot may also be, for example, an at least partially automated sweeping robot.
[0076] In particular, the vehicle can be stopped and / or completely switched off based on the value of the classification vector. If, for example, a risk to a living being, in particular a person, is derived based on the classification vector, a corresponding control is used to increase the operating safety of the corresponding vehicle. In the case of logistics, sweeping and / or mowing robots, accidents, in particular accidents with workers, pets and children, can also be prevented in this way.
[0077] According to another embodiment of the method, it is provided that a driving route is determined from a plurality of driving routes for an at least partially automated vehicle or for a driver assistance system as a function of at least one of the values of the classification vector. In the case of a driver assistance system, the determination involves a recommendation to the driver. That is, if the value of the classification vector indicates that an acoustic signal characterizing a dangerous situation has been identified, then, depending on the current traffic situation, a driving route can be determined which, for example, helps to form a free passage or, for example, can determine a driving route with a reduced speed.
[0078] A computer program is proposed, which comprises instructions which, when the program is executed by a computer, cause the computer to carry out the above method.
[0079] The computer program includes: a program code in any programming language; a computer program; a compiled version of the program code; firmware, by means of which the program code is implemented; or also a chip, the functionality of which represents the program code.
[0080] Furthermore, a machine-readable storage medium is proposed, which includes instructions, which, when executed by a computer, prompt the computer to implement the above method.
[0081] Furthermore, a machine-readable storage medium is proposed, on which a computer program is stored.
[0082] According to the present invention, a device is proposed, which is designed to implement the above method.
[0083] The device may be, in particular, a control unit, for example for an at least partially automated robot, in particular for an at least partially automated vehicle.
[0084] According to the present invention, a decision system for a driving route of a vehicle is proposed, wherein the decision system is designed to implement one of the above methods and determine a driving route from a plurality of driving routes of the vehicle according to the value of a classification vector. Such a decision system can be arranged in an at least partially automated vehicle and in a driver assistance system for use. BRIEF DESCRIPTION OF THE DRAWINGS
[0085] An embodiment of the present invention is shown with reference to FIGS. 1 to 2 and is further described below. In particular:
[0086] Figure 1a shows the order of frequency components with respect to time in the absence of a special signal;
[0087] Figure 1b shows the order of frequency components with respect to time with a follow-up signal;
[0088] Figure 1c shows the order of frequency components with respect to time in the case of a Wail signal;
[0089] Figure 2 A method for classifying an acoustic signal is shown. DETAILED DESCRIPTION
[0090] This exemplary embodiment uses three different digital audio data that follow one another in time and describe different acoustic signals as an example to illustrate how these are classified according to the method of the present invention.
[0091] according to Figure 2 In a first step S1, digital audio data that follow one another in time are transformed from the time domain into the frequency domain. This can be done, for example, by means of a short-time Fourier transform.
[0092] When calculating multiple frequency expressions for the progressive time interval of audio data that follow each other in time, a time window of 0.2 seconds is converted into a frequency domain with, for example, 2048 amplitudes from the beginning of approximately three seconds of digital audio data that follow each other in time by means of a short-time Fourier transform, thereby creating a frequency expression for the first time window. Then, the time window continues to move in time with 0.1 seconds in order to perform another short-time Fourier transform and form a second frequency expression. The steps are repeated until the algorithm has reached a terminal of approximately three seconds. This results in 28 frequency expressions for approximately three seconds.
[0093] The frequency representation is divided into its octaves, i.e., frequency ranges in which the end of the range is formed by twice the value of the frequency at the beginning of the range. Within each such octave, a determined number of frequency segments S2 is formed for each individual frequency representation. The determined number can be, for example, 12. Exemplarily, the frequency segments are formed uniformly and arranged side by side in the octave and include a subset of each frequency representation in the individual frequency representations.
[0094] For example, the values of 12 frequency segments are summed up S3 by the corresponding frequency segments of different octaves expressed by each single frequency, ie, the segments arranged at the same position within the octave. Thus, 12 added frequency segments are then formed, wherein the average value is respectively formed to calculate S4 the frequency component.
[0095] These 12 frequency components of the 28 frequency expressions are passed as input values to the classifier, for example as a 28×12 image, which generates an S5 classification vector based on this. The value of the classification vector classifies the signal characterizing the dangerous situation described by the digital audio data that follows each other in time with the help of the number of frequency components of the multiple frequency expressions and associates the classified classification vector with the corresponding value.
[0096] In this embodiment, the classifier has an artificial neural convolutional network that is designed and trained to classify the number of frequency components of a signal representing a dangerous situation described by temporally successive digital audio data expressed using multiple frequencies by generating the value of a classification vector.
[0097] Here, the artificial neural convolutional network is composed of a sequence of two blocks with convolutional layers.
[0098] The first block has the following layers:
[0099] Zero-padding layer: Rückskalierung of input with +1 in both image directions
[0100] Convolutional layers with core number N1 (e.g. N1=16) 3×3 and stride 1, ReLU activation function Zero-padding layer; degradation of input with +1 in both image directions
[0101] Convolutional layer with core number N1 (e.g. N1=16) 3×3 and stride 1, ReLU activation function
[0102] A max pooling layer of size 2×2 and stride 1
[0103] Lost Layer
[0104] Second block:
[0105] Zero padding layer; downgrading of input with +1 in both image directions
[0106] Convolutional layer with core number N1 (e.g. N1=16) 3×3 and stride 1, ReLU activation function
[0107] Zero padding layer; downgrading of input with +1 in both image directions
[0108] Convolutional layer with core number N1 (e.g. N1=16) 3×3 and stride 1, ReLU activation function
[0109] A max pooling layer of size 2×2 and stride 2
[0110] Lost Layer.
[0111] In order to complete the network model structure, a fully connected layer and a densely connected output layer are included with an activation function "Softmax". The dropout layer randomly suppresses some neurons in the neural network in order to reduce the possibility of overfitting. The last layer has an output variable corresponding to the number of categories of signals representing the category of dangerous situations.
[0112] Layer (Type) Original form parameter# Zero padding 2d1 (30,14,1) 0 Convolution2d1 (28,12,16) 160 Zero padding 2d2 (30,14,16) 0 Convolution 2d2 (28,12,16) 2320 Max Pooling 2d1 (14,6,16) 0 Lost 1 (14,6,16) 0 Zero padding 2d3 (16,8,16) 0 Convolution 2d3 (14,6,32) 4640 Zero fill 2d4 (16,8,32) 0 Convolution2d4 (14,6,32) 9248 Max Pooling 2d2 (7,3,32) 0 Lost 2 (7,3,32) 0 Flat 1 (672) 0 Dense 1 (64) 43072 Lost 3 (64) 0 Output Node (Number of categories) (64+1)×number of categories
[0113] Table 1 describes the layers in more detail.
[0114] The input is a 28×12×1 image or tensor data model
[0115] A suitable neural network according to the invention is trained by providing the number of frequency components of a plurality of frequency expressions as training data in the input layer and comparing the output data of the neural network with the desired classification. Subsequently, the parameters of the neural network are modulated until the agreement is sufficiently accurate (superwised lerning in English).
[0116] Figures 1a to 1c An example of an input value is shown, which is passed to a classifier. The abscissa is the time axis and the ordinate indicates 12 frequency components. The blackening of the small partial areas 10, 12 of the respective diagram is proportional to the height of the value of the frequency component.
[0117] Figure 1a shows the value of the frequency component when no special signal is detected in the audio data. Figure 1b In FIG. 1 , the follow-up signal is clearly visible due to the time-alternating frequency components having the highest amplitude of 10. Figure 1cThe Yelp special signal is shown, in which the frequency component with a maximum amplitude of 12 passes through different of the twelve frequency components in time succession. It is clear that the method is suitable for preprocessing special signals from traffic situations so that classification of the image is possible using different classifiers.
Claims
1. A method for classifying digital audio data that follow one another in time, the audio data describing acoustic signals that characterize a dangerous situation, the method comprising the following steps: calculating a plurality of frequency representations for progressively progressive time intervals of the audio data that follow each other in time (S1); forming a determined number of frequency segments (S2) for each octave of each single frequency expression, wherein the frequency segments comprise a subset of the single frequency expressions; Adding the corresponding frequency bins of the octaves expressed by each single frequency (S3); calculating the frequency components by forming an average value for each summed frequency bin in each single frequency representation (S4); A classification vector (S5) is generated with the aid of a classifier and the number of frequency components of multiple frequency expressions, wherein the classifier is designed to classify a signal characterizing a dangerous situation and described by corresponding digital audio data that follow one another in time with the aid of the corresponding number of frequency components of the multiple frequency expressions and to associate the classified classification vector with corresponding values.
2. The method according to claim 1, The number of frequency segments per octave for each individual frequency representation is arranged in an octave such that the frequency segments at least partially overlap one another.
3. The method according to claim 1 or 2, Digital audio data that follow one another in time are analyzed for the presence of fundamental tones and overtones, and frequencies in a frequency band around the fundamental tones and overtones in the audio data are attenuated before a plurality of frequency expressions are calculated for progressive time segments of the audio data that follow one another in time.
4. The method according to claim 1 or 2, The calculation of a plurality of frequency representations for progressive time intervals of the audio data which follow one another in time is carried out by means of a short-time Fourier transformation or a wavelet transformation.
5. The method according to claim 1 or 2, The frequency components of each single frequency expression are normalized before generating the classification vector.
6. The method according to claim 5, In this case, the frequency components of each individual frequency expression are normalized to the value one, or the frequency components of each individual frequency expression are normalized by means of a histogram equalization method.
7. The method according to claim 1 or 2, The classifier has an artificial neural feedforward network which is designed and trained to classify the number of frequency components of a signal characterizing a dangerous situation described by means of multiple frequencies and described using temporally successive digital audio data by generating values of a classification vector.
8. The method according to claim 1 or 2, The classifier has an artificial neural feedback network which is designed and trained to classify the number of frequency components of a signal representing a dangerous situation described by means of multiple frequencies and which is described using temporally successive digital audio data, by generating the value of a classification vector.
9. The method according to claim 1 or 2, The classifier has an artificial neural convolutional network which is designed and trained to classify the number of frequency components of a signal representing a dangerous situation described by means of temporally successive digital audio data expressed by means of a plurality of frequencies by generating the value of a classification vector.
10. The method according to claim 1 or 2, The classifier is designed to classify the signal representing a dangerous situation by means of a stepwise comparison of the number of frequency components of a plurality of frequency expressions with the number of templates and to associate a classification vector with different values depending on the result of the comparison.
11. The method according to claim 1 or 2, The classifier comprises a pre-classifier and a main classifier, and the pre-classifier is designed to identify a signal representing a dangerous situation described by digital audio data that follow each other in time by the number of frequency components expressed by multiple frequencies, and the main classifier is designed to classify the signal representing a dangerous situation described by digital audio data that follow each other in time by the value of a classification vector if the pre-classifier has identified the signal representing a dangerous situation in the digital audio data that follow each other in time.
12. The method according to claim 1 or 2, Therein, an actuation signal for actuating the at least partially automated vehicle and / or a warning signal for warning a vehicle occupant is emitted as a function of at least one of the values of the classification vector.
13. A device, The device is designed to carry out the method according to any one of claims 1 to 12.
14. A computer program product comprising a computer program, The computer program comprises instructions which, when the computer program is executed by a computer, cause the computer to carry out the method according to any one of claims 1 to 12 .
15. A machine-readable storage medium, A computer program is stored on the machine-readable storage medium. The computer program includes instructions that, when the computer program is executed by a computer, cause the computer to perform the method according to any one of claims 1 to 12 .
Citation Information
Patent Citations
Sound identification systems
CN102246228A
Method for identifying signs of faults occurred to electric tools
CN105004497A