Method for real-time analysis of a sound signal, and corresponding analysis device
The described sound analysis method addresses computational and privacy issues by using filtering and adaptive thresholding, enabling efficient and accurate detection of sound events in various environments without high computational demands or data transfer.
Patent Information
- Application Number
- FR2024004627
- Authority / Receiving Office
- FR · FR
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-02
- Publication Date
- 2025-11-07
AI Technical Summary
Existing real-time sound analysis methods struggle with high computational requirements, false positives in noisy environments, and privacy concerns due to the need for large databases and data transfer, making them unsuitable for low-power devices and environments with strict data protection laws.
A real-time sound analysis method involving high-pass and band-pass filtering, moving averages, and adaptive thresholding based on long-term and short-term averages to detect sound events efficiently, reducing computational demands and eliminating the need for pre-recorded databases.
The method effectively detects sound events with reduced false positives and computational requirements, suitable for low-power devices and compliant with privacy regulations, while maintaining accuracy in both quiet and noisy environments.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
Title of the invention: Method for real-time analysis of a sound signal, and corresponding analysis device. Field of the invention
[0001] The field of the invention is that of methods for real-time analysis of a sound signal.
[0002] In particular, the invention relates to such a method of analyzing a sound signal aimed at detecting, in this sound signal, particular sound events.
[0003] The invention also relates to detectors, or detection devices, implementing such a method of analyzing a sound signal allowing the detection of sound events. Previous art
[0004] Various methods for analyzing a sound signal in real time are known. For example, it is known to analyze the sound level of a signal, its frequency, etc. It is also known to analyze the sound level of a signal in order to detect a specific sound event corresponding to the exceeding of a threshold by this sound level.
[0005] Methods for identifying particular sounds in a complex sound signal are also known.
[0006] We know, for example, from the article by Hengel, Peter & Andringa, TC (2007). “Verbal aggression detection in complex social environments” (15-20. 10.1109 / AVSS.2007.4425279), of a real-time analysis method of a sound signal, including a calculation of the characteristics of a captured sound signal, in order to issue an alarm in the event of detection of a sound event.
[0007] In certain environments, such a method can effectively detect sound events corresponding to bursts of voices. However, it appears that such a method leads to the detection of a very large number of false positives when implemented in a noisy environment. Furthermore, it requires significant computing power to continuously analyze the signal characteristics that may be representative of such a sound event.
[0008] Such analyses can be implemented using computer-based machine learning methods, which in practice involve comparing samples of the sound signal to a library of sound samples containing identified sounds, in order to match the sound sample to the identified sound with the greatest similarity. Such methods generally require the use of computer systems such as artificial neural networks.
[0009] Implementing such methods requires significant computing power, substantial working memory, and a very large database of reference sounds. Consequently, it is difficult to implement them in a simple computer terminal such as a mobile phone or a detection device. On the contrary, implementing such a method often requires transferring the sound to be analyzed to a high-performance computer server capable of performing the analysis.
[0010] Such a transfer requires a high-capacity connection. Moreover, the transfer of a sound recording to a remote server may, in many cases, be prohibited by privacy and personal data protection laws.
[0011] There is therefore a need to provide other methods for real-time analysis of a sound signal and, in particular, for the detection of a sound that can be considered an anomaly, or a sound event that stands out in a sound environment.
[0012] A particular objective of the invention is to provide such a method for real-time analysis of a sound signal enabling the reliable detection of sound events that may correspond to bursts of voices.
[0013] Another objective of the invention is to provide such a method of real-time analysis of a sound signal enabling the efficient detection of such sound events, including in a noisy sound environment, by limiting the number of erroneous detections, known as "false positives".
[0014] In particular, an objective of the invention is to provide such a method which requires relatively low computing power and memory capacity.
[0015] Another objective is to provide such a detection method which does not require the use of pre-recorded sound databases.
[0016] Another objective is to provide such a detection method which can be implemented without storing or transmitting sound recordings which could be considered personal data. Description of the invention
[0017] These objectives, as well as others which will become clearer later, are achieved using a real-time analysis method for a sound signal, comprising the following steps: - a step of capturing a sound signal by at least one sensor; - a high-pass filtering stage of the captured sound signal, cutting off at least the frequencies below 100 Hz; - a step of rectifying the captured sound signal; - a step of calculating a first moving average of the filtered and rectified signal over a first sliding time period, between 1 ms and 20 ms; - a step of calculating a second moving average of the rectified signal over a second sliding time period of a duration between 100 ms and 10 s, and greater than 10 times the duration of the first sliding time period; - a step of comparing the first moving average with a threshold; - a step involving the emission of a signal representative of a sound event, in case of exceeding the threshold by the first moving average; the threshold being variable depending on the value of the second moving average (LTA), such that: - when the value of the second moving average is less than a predefined value V, between 50 dB and 75 dB, the threshold is between 60 dB and 85 dB, and greater than V + 6 dB; - when the value of the second moving average is greater than the predefined value V, the threshold is between the value of the second moving average plus 6 dB and the value of the second moving average (LTA) plus 14 dB.
[0018] Such a method advantageously allows for the efficient detection of sound events in a noisy environment. It proves to be very effective both in a quiet environment, where the relatively high detection threshold helps to avoid false alarms, and in a noisy environment, where the threshold adapts to the sound level.
[0019] Advantageously, the high-pass filtering cuts off at least the frequencies below 500 Hz.
[0020] Such filtering makes it possible to avoid taking into account, in the detection, low frequency sounds which are very present on public roads, such as the sounds emitted by vehicles.
[0021] According to a preferred embodiment, the filtering is a bandpass filter cutting off at least the frequencies above 8 kHz.
[0022] Such filtering makes it possible to avoid taking into account, in the detection, high frequency sounds which are not very relevant for the detection of sound events.
[0023] According to preferred embodiments, the bandpass filtering cuts off at least the frequencies above 4 kHz, and preferably at least the frequencies above 3 kHz.
[0024] Such filtering, cutting off frequencies above 4 kHz, advantageously allows detection to be focused on the frequencies of the human voice. The filtering cutting off the Frequencies above 4 kHz advantageously limit the consideration of sounds generated by babies crying in the detection of sound events.
[0025] According to an advantageous embodiment, the first moving average and the second moving average (LTA) are quadratic means of the filtered signal.
[0026] In such a case, the calculation of the root mean square makes it possible to perform in a single calculation the step of rectifying the sound signal and the step of calculating the average.
[0027] Preferably, the first sliding time period has a duration of between 2 and 5 ms.
[0028] Preferably, the value V is between 60 dB and 70 dB.
[0029] It has been observed that such a choice makes it possible to limit the number of false alarms.
[0030] According to an advantageous embodiment, when the value of the second average mobile is less than the predefined value V, the threshold S is equal to a constant value C, between 60 dB and 85 dB, and greater than V + 6 dB.
[0031] Preferably, when the value of said second moving average is less than said predefined value V, said threshold S is between 70 dB and 80 dB.
[0032] According to an advantageous embodiment, when the value of the second moving average is greater than the value V, the threshold is greater than the value of the second moving average plus 10 dB.
[0033] According to an advantageous embodiment, the value of the threshold (S) is defined by the relation S = max (C, LTA + k), the value of C being a constant between 70 dB and 80 dB, and the value of k being a constant between 6 dB and 14 dB.
[0034] The present invention also relates to a real-time analysis device for a sound signal, comprising a sensor capable of capturing a sound signal, calculation means and means for emitting a signal, this device being configured for the implementation of the analysis process described above.
[0035] List of figures
[0036] The invention will be better understood upon reading the following description of preferred embodiments, given by way of simple figurative and non-limiting example, and accompanied by the figures, among which: - Fig. 1 is a schematic representation of the different stages of a sound signal analysis process according to one embodiment of the invention. - Figure 2 is a graph showing the representative curve of a signal sound that can be captured by a sensor such as a microphone. - Fig. 3 is a graph showing the curve representing the absolute value of the sound signal represented by the curve in Fig. 2, constituting the rectified signal. - The [Fig.4] is a graph superimposing a curve representing a first moving average of the rectified signal represented by the curve of [Fig.3], a curve representing a second moving average of the rectified signal represented by the curve of [Fig.3], and a curve representing a variable threshold, according to an embodiment of the invention, as a function of the second moving average. - Fig. 5 is a graph representing the curves of Fig. 4, represented on a logarithmic scale. Figure 6 is a graph showing the variation of the threshold as a function of the second moving average, according to one embodiment of the invention. Detailed description of embodiments of the invention
[0037] Fig. 1 schematically represents the steps of a sound signal analysis process according to one embodiment of the invention.
[0038] These steps are preferably implemented sequentially, repeatedly at a high rate, until a sound event is identified. Typically, they are implemented by a processor analyzing a time sample of digital data from a digital microphone or an analog-to-digital converter, before analyzing the next time sample. Alternatively, these steps can also be implemented continuously by an electronic circuit analyzing analog data from a microphone.
[0039] In this analysis process, a first step 11 corresponds to the capture of a sound signal. This sound signal may correspond to the signal captured by one or more sensors, such as microphones, which may optionally undergo pre-processing such as the removal of its continuous component and digitization.
[0040] The graph in [Fig. 2] shows the curve 21 representing such a sound signal captured by a sensor such as a microphone, which consists of oscillations around a neutral value 20. This sound signal may, for example, correspond to sounds captured in a public place. In [Fig. 2], its amplitude is normalized so that the maximum and minimum values of the signal correspond respectively to 1 and -1.
[0041] This sound signal, captured by one or more microphones, advantageously undergoes a real-time filtering step 12. This filtering step advantageously removes all the frequency components of the sound signal that do not need to be analyzed.
[0042] Thus, the sound signal undergoes, according to the invention, a high-pass filtering cutting off at least the frequencies below 100 Hz and, preferably, at least the frequencies below 500 Hz.
[0043] Such filtering often makes it possible to significantly reduce the volume of the filtered signal compared to the volume of the captured signal. Indeed, in a noisy environment, a large part of the noise consists of low-frequency sounds. This is particularly true of a large proportion of engine noise and vehicle rolling noise.
[0044] When the analysis method according to the invention is implemented on the sound captured in a vehicle, such as a bus, the application of a high-pass filter cutting off frequencies below 500 Hz thus makes it possible, in certain cases, to obtain a filtered sound signal in which the sound volume is 15 dB lower than the sound volume of the captured sound signal.
[0045] In a preferred embodiment, the audio signal undergoes bandpass filtering that also cuts off frequencies above 8 kHz and, preferably, frequencies above 4 kHz. Cutting off the high frequencies also significantly reduces the volume of the filtered signal compared to the volume of the captured signal.
[0046] The cutoff frequencies used for filtering are preferably chosen according to the nature of the sound events to be detected. When the sound events to be detected are caused by human voices, for example, bursts of voice, a bandpass filter cutting off frequencies below 500 Hz and frequencies above 4 kHz is preferable. Indeed, sound signals from human voices do not contain frequencies above 4 kHz. Such frequencies are therefore useless for detecting bursts of voice.
[0047] If it is desired to prevent the crying of babies from being detected as sound events by the analysis method of the invention, it is advantageous to use a bandpass filter cutting off frequencies above 3 kHz and, for example, frequencies below 500 Hz.
[0048] When, on the contrary, the sound events that one wishes to detect are machine noises, for example the noise of a drill, one can advantageously use a bandpass filter cutting off frequencies below 1 kHz and above 8 kHz.
[0049] More generally, in variants of this embodiment, the cut-off thresholds can be set to different values to adapt to the sound sources causing the sound events that the device must detect.
[0050] This signal, filtered during the filtering step 12, is then processed during a rectification step 13 to obtain a rectified signal. For this purpose, it is possible, for example, to calculate the absolute value of the filtered signal. According to another method of In practice, it is also possible to calculate the square of the filtered signal value. This square, being always positive, is a rectified signal representative of the filtered signal.
[0051] The graph in [Fig. 3] thus shows the curve 22 representing the absolute value of the sound signal represented by the curve 21 of [Fig. 2]. This rectified signal, representing the sound pressure of the filtered signal, is constantly positive with respect to the neutral value 20, and it is possible to calculate average values, commonly called "envelopes", representing the sound volume of the filtered sound signal.
[0052] In the embodiment shown, the captured sound signal first undergoes filtering and then rectification. This order of steps is advantageous, as filtering is often more effective on the raw sound signal than on the rectified sound signal. However, in an alternative embodiment, it is also possible for the sound signal to be filtered based on the rectified signal.
[0053] Sound events, as defined in this description, correspond to an unexpected evolution of a sound signal within a given sound environment. These sound events correspond to a physical reality, but the definition of what constitutes a sound event can, of course, vary depending on the circumstances.
[0054] By way of example, the present invention is particularly well suited for enabling the detection of sound events such as bursts of voices, including in noisy environments. It is normal for the ambient sound of an environment where people are present to include human voices. These voices can sometimes be quite loud, for example, when the environment is noisy, forcing people to speak louder to be heard. However, these voices do not necessarily constitute unexpected sound events that warrant being reported.
[0055] On the other hand, in such environments, one can sometimes hear outbursts of voices corresponding to sudden, loud noises, generally expressing strong emotion, anger, and / or aggression. Detecting such audible events, which are often linked to assaults, is particularly useful. Their detection can, for example, allow for the focus of attention from surveillance services and, potentially, the intervention of security services.
[0056] In other situations, the sound events that one seeks to detect may be of a different nature. For example, it could be an event such as a detonation, or a sound caused by a particular machine, such as a drill...
[0057] To perform the detection of sound events in the sound signal, a first moving average is calculated in real time, during step 14, over a first sliding time period, of the filtered and rectified sound signal. This The first moving average is subsequently called the "instantaneous average" of the filtered and rectified signal, or "STA" (acronym for the English expression "Short Terni Average"). It is preferably representative of the envelope of the rectified signal, corresponding to the instantaneous variations in the amplitude of the sound pressure. For this, this instantaneous average STA is calculated over a first sliding time period Tl, or first sliding time window, of a given relatively short duration.
[0058] The choice of the duration of this first period of time Tl can vary depending on the type of sound events that one seeks to detect.
[0059] When the sound events to be detected are caused by human voices, for example, bursts of voice, this first time period Tl is preferably greater than 1 ms, in order to prevent the instantaneous average STA from being representative of very brief sounds, such as a brief impact on a microphone. This first time period Tl is also preferably less than 20 ms, so that the instantaneous average STA is representative of brief sounds constituting speech, such as an interjection.
[0060] This first time period Tl is therefore preferably between 1 ms and 20 ms. Preferably, it is between 2 ms and 10 ms and, even more preferably, between 2 ms and 5 ms. Such a range of values has indeed proven particularly effective for detecting, in a noisy environment, a sound event corresponding to a burst of voice.
[0061] Thus, in the embodiment of the invention shown in the figures, during step 14 of the process according to the embodiment of the invention shown, an average value STA of the filtered and rectified sound signal is calculated in real time over a first time period Tl of 4 ms that has just elapsed. This calculation of the average value STA is preferably performed in real time, the average value at each instant being calculated over the first time period Tl preceding that instant.
[0062] Curve 23, in the graph of [Fig. 4], represents, on a linear scale, the evolution over time of the value of the instantaneous arithmetic mean STA, representative of the envelope of the filtered and rectified sound signal represented by curve 22 in [Fig. 3]. The graph of [Fig. 5] represents the same curve 24 on a logarithmic scale.
[0063] To perform the detection of sound events in the sound signal, a second moving average is also calculated in real time during step 15, over a second sliding time period of the filtered and rectified sound signal. This second moving average is subsequently called the "long average" of the filtered and rectified sound signal, or "LTA" (acronym for the English expression "Long"). Term Average). It advantageously constitutes a representative variable of the sound environment over a second sliding time period T2, or second time window, of a given relatively long duration, of the sound signal to be analyzed.
[0064] This second time period T2 is chosen to be significantly longer than the first time period T1. In general, the duration of this second period T2 is thus greater than 10 times the duration of the first period T1.
[0065] Preferably, the duration of this second time period T2 is between 100 ms and 10 s.
[0066] Thus, during a step 15 of the process according to the embodiment of the invention shown, a long average (LTA) of the processed signal, and therefore of the absolute value of the filtered sound signal, is calculated in real time over a second time period T2, which has just elapsed, with a duration of 1 s. Such a time period allows the average value to be representative of a sound environment at a given moment. It is, however, possible to choose this second time period T2 to be between 100 ms and 10 s.
[0067] This calculation of the long average LTA is preferably carried out in real time, the value of the average at each instant being calculated over the time period T2 preceding that instant.
[0068] Curve 24, in the graph of [Fig. 4], represents, on a linear scale, the evolution over time of the value of this long average LTA calculated on the basis of the rectified filtered sound signal represented by curve 22 in [Fig. 3]. The graph of [Fig. 5] represents the same curve 24 on a logarithmic scale.
[0069] As shown in curves 23 and 24, the long average LTA is relatively insensitive to rapid variations in the amplitude of the sound signal. On the contrary, the instantaneous average STA follows the envelope of the sound signal.
[0070] The calculation of the instantaneous average (STA) and long-term average (LTA) values can be carried out in several ways known to those skilled in the art. These averages can, for example, be arithmetic means, root mean squares, or harmonic means. In particular embodiments, these averages can be weighted, for example, to give more weight to more recent values in the time period T1 or T2.
[0071] In an advantageous embodiment, a root mean square (RMS) method is used. Such a method gives greater weight to the highest sound values, as these sounds are particularly important for detecting sound events. Furthermore, the use of the RMS method is consistent with the definition of sound level in decibels.
[0072] It should be noted that calculating the root mean square (RMS) involves squaring the sound signal. Calculating such a RMS on the filtered sound signal, but unrectified, allows the rectification step of the sound signal and the calculation of the average to be carried out at the same time.
[0073] In another particularly advantageous embodiment, the signal rectification step is performed by calculating the square of the filtered signal value, and the averaging step is performed by calculating an arithmetic mean. Such an arithmetic mean, based on previously squared values, is representative of a root mean square.
[0074] Thus, if the filtered sound signal is a digital signal, the instantaneous average STA, or respectively the long average LTA, can correspond to the arithmetic, or preferably quadratic, mean of the absolute values of this filtered sound signal during the first time period T1, or respectively the second time period T2. The instantaneous average STA, or respectively the long average LTA, can also be obtained by a low-pass filter of the filtered and rectified signal, whose time constant corresponds to the time period T1, or respectively to the second time period T2.
[0075] It should be noted that the implementation of the invention does not necessarily assume that the instantaneous average STA and the long average LTA are calculated by the same calculation methods.
[0076] It is possible to detect in real time a sound event corresponding to a sudden and abrupt increase in the amplitude of the sound signal by analyzing the evolution of the instantaneous average STA. To identify such sudden and abrupt increases in the amplitude of the sound signal, representative of the sound events to be detected, the instantaneous average STA is compared in real time, during a step 17, to a threshold S. It is thus considered that if the value of the instantaneous average STA, representative of the envelope of the sound signal STA, exceeds the threshold S, it means that the sound environment has undergone an unpredictable evolution characteristic of a sound event. In such a case, the analysis process is interrupted after a step 18 of emitting a signal to detect a sound event within the sound signal.
[0077] According to the invention, the threshold S, which the instantaneous average STA must exceed to trigger step 18 of emission of a signal for detecting a sound event, is not a fixed threshold, but has a variable value which is itself determined, during a step 16, as a function of the long average LTA.
[0078] Fig. 6 thus shows the curve of the evolution of the threshold S as a function of the long average LTA, in a preferred embodiment of the invention.
[0079] As this figure shows, below a given value V of LTA, which is predefined and between 50 dB and 75 dB, the threshold S is between 60 dB and 85 dB, and greater than V + 6 dB. On [Fig.6], the hatched area 61, for LTA values less than V = 64 dB, represents the values that the threshold S can take.
[0080] Preferably, below this value V, the threshold S is equal to a constant value C, between 60 dB and 85 dB, and greater than V + 6 dB.
[0081] Thus, in the embodiment shown, the threshold S is equal to a constant value C = 74 dB, as long as the value of LTA is less than V = 64 dB. This value is represented by the portion of curve 62 on [Fig.6].
[0082] In Figures 4 and 5, curve 25 represents the threshold S, respectively according to a linear scale and a logarithmic scale. When the LTA value remains below a value V = 64 dB, that is, before point 241 marked on curve 24 of LTA, the threshold S remains constant and equal to C = 74 dB, which corresponds to the straight segment 251 of curve 25, before point 252.
[0083] Thus, when the value of LTA is minimal or low, which corresponds to a situation of silence or very low ambient noise, the threshold S that must be crossed by STA for a sound event to be detected is constant and relatively high.
[0084] Such a high threshold advantageously prevents even the slightest sound, however low in intensity, from being detected as an anomaly in the sound environment in a quiet or near-quiet environment. This therefore avoids many false alarms. On the other hand, all STA exceedances of the 74 dB threshold are considered as sound events.
[0085] This constant value C is preferably chosen at the time of installation of the detection device, depending on the distance between this detection device and the foreseeable noise sources. For example, it may advantageously be possible to set this constant value C at 80 dB if this distance is 1 m, and to decrease it by 6 dB each time this distance doubles. This constant value C could thus be 74 dB if the distance is 2 m, 68 dB if the distance is 4 m, and so on.
[0086] When the value of LTA is greater than the predefined value V, the threshold S has a value that increases with LTA and is greater than LTA. Advantageously, when the value of LTA is greater than the predefined value V, the threshold S is defined by the relation S = LTA + k, with 6 dB < k < 14 dB
[0087] On [Fig.6], the hatched area 63, for LTA values greater than V = 64 dB, represents the values that the threshold S can take.
[0088] Thus, in the embodiment shown, when the value of LTA is greater than V = 64 dB, the threshold S is equal to LTA + k, with k = 10 dB. This value is represented by the portion of curve 64 on [Fig. 6].
[0089] Thus, in figures 4 and 5, when the value of LTA becomes greater than a value V = 64 dB, that is after the point 241 marked on the curve 24 of LTA, the threshold S represented by the segment 253 of the curve 25 is equal to LTA + 10 dB.
[0090] Such a definition of the threshold S allows, when the sound environment is considered noisy (the value of LTA being greater than 64 dB), that the exceedances by STA of a value of LTA + 10 dB are considered as sound events of the sound signal.
[0091] Thus, when the LTA value is high, corresponding to a noisy sound environment, the predetermined threshold S that STA must cross for a sound event to be detected is higher than LTA but relatively close to LTA. This threshold close to LTA thus advantageously allows the detection of a discontinuity in the sound signal when a sound event exceeds the average noise level, even if this exceedance is, proportionally, relatively small.
[0092] In the embodiment shown, we can therefore consider that there is a discontinuity in the sound signal when STA > max (C, LTA + ki), with ki a constant which, in the embodiment shown, is equal to 10 dB.
[0093] In the situation represented by the graphs in Figures 4 and 5, such an overshoot occurs eight times, at the time when the value of LTA forms peaks 331 to 338.
[0094] The various values defined above are expressed in dB, which corresponds to the logarithmic scale commonly used to measure sound levels. Figure 5 represents the values according to this scale. Of course, these same values can also be expressed linearly. Figures 1 to 4 represent them according to this linear scale.
[0095] According to a linear scale, in the embodiment shown above, there can be a discontinuity in the sound signal when STA > max (10e, k2 x LTA), with k2 a constant which, in the embodiment shown, is approximately equal to 3.
[0096] When, during step 17, the instantaneous average STA is measured as being below the threshold S, the process according to the invention continues without any identification of a sound event.
[0097] When, on the contrary, the instantaneous average STA is measured as greater than the threshold S, a sound event is identified, during a step 18. This step may, for example, correspond to the emission of a computer signal indicating the sound event, or to the triggering of an alarm.
[0098] In a preferred embodiment, the first identification of a sound event interrupts the real-time analysis process according to the invention. Another analysis process can then be implemented, for example to more precisely identify the nature of the sound event, before the real-time analysis process according to the invention resumes.
[0099] In such a case, in the situation represented by the graphs in figures 4 and 5, the identification process would be interrupted as soon as the peak 231 of the value of STA.
[0100] In other embodiments, it is also possible for the real-time analysis process according to the invention to continue despite the identification of a sound event. In this case, in the situation represented by the graphs in Figures 4 and 5, it can identify a sound event at the time of each peak 231 to 238 of the STA value.
[0101] In other embodiments, it is also possible for the real-time analysis process to continue despite the identification of a sound event, in a specific version. For example, the identification of a sound event may cause the threshold value S to be blocked for a limited time, for example, 3 seconds. Such a blockage makes it possible to detect a succession of sound events without the rise in the LTA value, due to the first sound events, causing a rise in the threshold S that would make it difficult to detect subsequent events.
Claims
Demands
1. A method for real-time analysis of a sound signal, characterized in that it comprises the following steps: - a step (11) of capturing a sound signal by at least one sensor; - a step (12) of high-pass filtering the captured sound signal, cutting off at least the frequencies below 100 Hz; - a step (13) of rectifying the captured sound signal; - a step (14) of calculating a first moving average (STA) of the filtered and rectified signal over a first sliding time period, between 1 ms and 20 ms; - a step (15) of calculating a second moving average (LTA) of the rectified signal over a second sliding time period of a duration between 100 ms and 10 s, and greater than 10 times the duration of said first sliding time period; - a step (17) of comparing the first moving average (STA) with a threshold (S);- a step (18) of emitting a signal representative of a sound event, in the event of exceeding said threshold (S) by said first moving average (STA); said threshold (S) being variable according to the value of said second moving average (LTA), such that: - when the value of said second moving average (LTA) is less than a predefined value V, between 50 dB and 75 dB, said threshold (S) is between 60 dB and 85 dB, and greater than V + 6 dB; - when the value of said second moving average (LTA) is greater than said predefined value V, said threshold (S) is between the value of said second moving average (LTA) plus 6 dB and the value of said second moving average (LTA) plus 14 dB.
2. An analysis method according to the preceding claim, characterized in that said high-pass filtering cuts off at least the frequencies below 500 Hz.
3. An analysis method according to any one of the preceding claims, characterized in that said filtering is a bandpass filtering that cuts off at least the frequencies above 8 kHz.
4. An analysis method according to the preceding claim, characterized in that said bandpass filtering cuts off at least frequencies above 4 kHz, and preferably at least frequencies above 3 kHz.
5. An analysis method according to any one of the preceding claims, characterized in that said first moving average (STA) and said second moving average (LTA) are quadratic means of said filtered signal.
6. Analysis method according to any one of the preceding claims, characterized in that said first sliding time period has a duration of between 2 and 5 ms.
7. Analytical method according to any one of the preceding claims, characterized in that said value V is between 60 dB and 70 dB.
8. Analysis method according to any one of the preceding claims, characterized in that, when the value of said second moving average (LTA) is less than said predefined value V, the threshold S is equal to a constant value C, between 60 dB and 85 dB, and greater than V + 6 dB.
9. Analysis method according to any one of the preceding claims, characterized in that, when the value of said second moving average (LTA) is less than said predefined value V, said threshold S is between 70 dB and 80 dB.
10. Analysis method according to any one of the preceding claims, characterized in that, when the value of said second moving average (LTA) is greater than said value V, said threshold (S) is greater than the value of said second moving average (LTA) plus 10 dB.
11. Analytical method according to any one of the preceding claims, characterized in that the value of said threshold (S) is defined by the relation S = max (C, LTA + k), the value of C being a constant between 70 dB and 80 dB, and the value of k being a constant between 6 dB and 14 dB.
12. A device for real-time analysis of a sound signal, comprising a sensor capable of capturing a sound signal, computing means and means of emitting a signal, characterized in that it is configured for the implementation of the analysis method according to any one of the preceding claims.
Citation Information
Patent Citations
Method for real-time processing of a sound signal and device for capturing a sound signal
EP4312215A1
Abuse Alert System by Analyzing Sound
US20210005069A1