Method and system for tamper detection of a microphone device
By sampling the audio signal of the microphone device in the low-frequency band and analyzing it with computer algorithms, tampering of the microphone device can be detected, which solves the problem of unreliable detection of microphone tampering in the existing technology and improves the reliability and effectiveness of the audio monitoring system.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- AXIS
- Filing Date
- 2025-12-18
- Publication Date
- 2026-06-23
AI Technical Summary
Existing technologies struggle to reliably detect microphone device tampering, especially sound blocking or muting, without manual listening or physical inspection, which affects the reliability and integrity of audio monitoring.
By sampling the audio signal from the microphone device in the low-frequency band, detecting the absolute amplitude of the audio signal decreasing exponentially over a certain period of time, and using computer algorithms to analyze the characteristic events in the audio signal, a tamper detection signal is output.
This technology enables reliable detection of microphone device tampering without human intervention, improving the reliability and effectiveness of the audio monitoring system, reducing the risk of false alarms, and ensuring automatic tampering detection of audio signals.
Smart Images

Figure CN122269207A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the detection of tampering with the microphone of a microphone device. More specifically, this disclosure relates to the detection of tampering that prevents the microphone from accurately capturing sound. Background Technology
[0002] Devices that include microphones typically consist of a microphone housed within a cavity that opens to the outside of the device. However, by covering or sealing the opening of the channel, sound can then be prevented from reaching the microphone within the channel. If the seal is too tight, sound will not be able to enter, thus rendering the microphone inoperable.
[0003] When a microphone is used as an audio sensor, the sound is analyzed by algorithms, making it difficult to determine whether the microphone has been tampered with without physically inspecting the device.
[0004] In this situation, the entire microphone function may be disabled without detection, compromising the integrity of the collected audio data.
[0005] This underscores the need for more reliable methods to detect tampering with devices, including microphones, ensuring that any interference, such as muting or blocking, can be identified and addressed. Therefore, improvements in tamper detection are required to maintain continuous and reliable audio monitoring. Summary of the Invention
[0006] The purpose of this disclosure is to enable reliable detection of tampering with microphone devices.
[0007] Another objective is to improve the reliability and integrity of audio monitoring systems without the need for human intervention, such as manual listening or physical inspection.
[0008] A further objective is to facilitate the automatic analysis of audio signals, allowing for accurate identification of tampering while minimizing false alarms.
[0009] To achieve at least one of the above objectives, and others as will become apparent from the following description, a method having the features defined in claim 1 is provided according to the invention. Preferred embodiments will become apparent from the dependent claims.
[0010] More specifically, according to a first aspect of the invention, a computer-implemented method for tamper detection of a microphone device is provided, the microphone device including a microphone disposed in a cavity open to the surrounding environment of the microphone device via a microphone hole, the method comprising: receiving an audio signal from the microphone; sampling the audio signal for a low frequency band; detecting an event, the event including an exponentially decreasing absolute amplitude of the sampled audio signal over a period of at least one second; and outputting a tamper detection signal when the event is detected.
[0011] Therefore, this method can reliably detect tampering with microphone devices, thereby improving the reliability of audio monitoring systems.
[0012] This ensures that any interference, such as microphone mute or blockage, can be quickly identified and resolved.
[0013] Generally, the method enables the detection of sound blockage. Therefore, it allows for the identification of when a microphone is intentionally blocked, ensuring that any attempt to disable the microphone by obstructing sound is reliably detected.
[0014] Furthermore, the method enables tampering detection without human intervention. For example, individuals can identify tampering without listening to the audio recorded by the microphone device or physically inspecting the microphone device. This eliminates the need for manual work in identifying microphone device disabling or tampering. Thus, the method enables more efficient and less labor-intensive identification of microphone device tampering.
[0015] In other words, the method enables the algorithm to analyze audio signals while still allowing for reliable identification of tampering. Therefore, it can provide reliable, automated tampering detection and analysis of audio signals.
[0016] Furthermore, the method enables improved tamper detection without requiring any changes to the microphone device or microphone hardware.
[0017] Furthermore, the method mitigates the risk of false alarm detection. By detecting characteristic events (including those with an exponentially decreasing absolute magnitude of the audio signal sampled over a period of at least one second), the method can distinguish between explicit tampering and other anomalies.
[0018] Therefore, this method provides a robust and effective solution for detecting tampering with microphone devices. By enabling reliable detection of sound blocking and tampering without human intervention, the method enhances the reliability and effectiveness of audio monitoring systems.
[0019] The term "exponentially decreasing absolute amplitude" can refer to a specific pattern of amplitude reduction in a sampled audio signal. In this context, "exponential" means that the rate of decrease is proportional to the current amplitude, resulting in a relatively rapid initial drop that slows down over time. Mathematically, this can be described using an exponential function, where the derivative of the amplitude with respect to time is linearly proportional to the amplitude itself, resulting in a smooth and continuous decay to the zero state of the sampled audio signal.
[0020] The absolute magnitude that decreases exponentially can be called the ramp in the sampled audio signal. The ramp can approach the zero state of the sampled audio signal over time.
[0021] A ramp can be defined by (absolute) exponential decay, where the amplitude decreases at a rate proportional to its current value, resulting in an initial decrease that gradually slows as it approaches zero. Including the absolute amplitude (i.e., the magnitude of the amplitude) ensures that both positive and negative values of the audio signal are considered. The term "tampering," as used herein, can refer to any attempt to disable or interfere with the ability of a microphone device to record audio from its surroundings. Tampering can involve placing or attaching an object such as chewing gum, cloth, tape, a hand, or any other item over the microphone hole in the cavity in which the microphone is disposed. Tampering can be any act intended to cover or block the microphone by covering the microphone hole. Furthermore, tampering can include any method that impairs the function of the microphone, such as applying a substance that silences sound.
[0022] The system can output or present tamper detection signals to alert users or supervisors of audio monitoring systems, including those with microphones. Tamper detection signals can be output, for example, in the form of, but not limited to, audible alarms, digital messages such as error messages, text messages displayed on a screen, visual alarms such as flashlights, or digital notifications such as alarms in software applications.
[0023] The microphone device can be a device that includes a microphone (e.g., a camera, surveillance camera, mobile phone, wearable device, etc.). In particular, the microphone device can be any suitable device that includes a microphone housed or arranged in a cavity that is open to the surrounding environment of the microphone device through a microphone hole.
[0024] A microphone hole can be referred to as the opening of a cavity. Generally, a cavity can be a microphone compartment or a microphone channel configured to house a microphone.
[0025] The microphone hole can be the only opening of the cavity. Therefore, there may be no connection, hole, or opening between the cavity and the interior of the microphone device. In other words, the cavity can form a closed compartment that is only open to the surrounding environment of the microphone device.
[0026] As used herein, the term "audio signal" can refer to the electrical representation of sound. This includes digital signals, which are discrete representations of sound produced by sampling analog signals at regular intervals. When a microphone captures sound from its surroundings or environment, it can generate an audio signal, converting acoustic energy into an electrical signal that can be processed, transmitted, or recorded. An audio signal can contain information about the amplitude, frequency, and / or phase of the sound. An audio signal can be represented, for example, by the movement (i.e., displacement) of the microphone diaphragm. An audio signal can be represented as a graph of amplitude changing over time; for example, the movement of the microphone diaphragm corresponds to the amplitude of the audio signal.
[0027] It should be further understood that the term "absolute amplitude" refers to the absolute value of the amplitude.
[0028] Generally, a microphone is used as a transducer to convert sound (pressure) into electrical current. It consists of a diaphragm that vibrates in response to sound waves. The electrical signal received from the microphone corresponds to the position of the diaphragm at each moment. When an object presses against the microphone aperture, the movement of the diaphragm is impeded, resulting in a reduction or complete cessation of the audio signal.
[0029] When an object presses against the microphone aperture, a pressure change occurs within the cavity (e.g., an increase in pressure). This pressure change causes the microphone diaphragm to move in a distinctive manner. The resulting movement of the microphone diaphragm (i.e., high absolute amplitude with low frequency) differs from the diaphragm movement caused by normal or conventional sound.
[0030] Specifically, when something presses against the microphone aperture, the microphone diaphragm displaces with a relatively large and slow movement. The diaphragm thus exhibits a DC offset, meaning it shifts from its neutral position due to changes in pressure within the cavity. However, the diaphragm will eventually return to the DC position (i.e., the center position). Note that this return is relatively slow and follows exponential behavior.
[0031] The method described herein relates to detecting events involving the aforementioned behavior of audio signals (i.e., the behavior of the microphone diaphragm). This enables reliable detection of tampering with the microphone device (e.g., corresponding to covering the microphone aperture).
[0032] The process of sampling an audio signal for the low-frequency band can include decimating the audio signal to reduce the sampling rate to approximately 200 Hz. This decimation process can be performed in multiple stages. Initially, the audio signal can be downsampled by a factor of 2, effectively halving the sampling rate. This step can then be followed by another downsampling stage of a factor of 3, reducing the sampling rate to one-third of its previous value. These stages can be performed according to established signal processing techniques. To maintain the integrity of the audio signal, additional low-noise low-pass filters can be applied between the final decimation stages. This filter can be of the scaling paradigm (SNF) type, which can be designed to minimize noise and maintain the quality of the audio signal during the decimation process.
[0033] Generally, the low-frequency band can include frequencies below 500 Hz. However, it should be recognized that the low-frequency band can, for example, include frequencies below 1 kHz, preferably below 500 Hz, or more preferably below 250 Hz.
[0034] Low-frequency bands can be implemented using low-pass filters. Low-pass filters can be designed to detect and process frequencies within a specified range, such as approximately 0.1 Hz to 500 Hz. In the example, a low-pass filter could sample frequencies below 250 Hz or below 200 Hz.
[0035] Therefore, low-frequency audio signals can pass through, while high-frequency noise can be attenuated. This can be beneficial for the detection of events in the audio signal.
[0036] Furthermore, by improving the signal-to-noise ratio, low-pass filters can make audio signals clearer and more reliable for further processing or analysis to detect events.
[0037] In a sense, when an event occurs at a low frequency, the audio signal from the microphone can be downsampled.
[0038] The absolute magnitude of the event decreasing exponentially during the period can be between 1 and 20 seconds.
[0039] In other words, it may take 1 to 20 seconds for the microphone diaphragm to reach or return to DC. Therefore, when the absolute amplitude decreases, the amplitude of the sampled audio signal approaches DC (i.e., corresponding to the diaphragm's resting or zero state).
[0040] In the example, the time period can be from 2 to 10 seconds. However, it should be recognized that time periods with an exponentially decreasing absolute magnitude can be characteristic of different microphone types or different microphone devices. The time period can, for example, depend on the size of the microphone, cavity, and / or microphone aperture. Therefore, for a specific type of microphone device, the time period can be a known or predetermined value.
[0041] This can provide enhanced reliability in event detection. For example, it can help to further eliminate false alarms.
[0042] In the example, the event may include the peak absolute amplitude of the sampled audio signal, followed by an exponentially decreasing absolute amplitude of the sampled audio signal over a period of at least one second.
[0043] The event detection step may further include: detecting the absolute amplitude peak in the sampled audio signal before the absolute amplitude of the sampled audio signal decreases exponentially.
[0044] The peak absolute amplitude can be detected before the absolute amplitude decreases exponentially, for example, within 3 or 5 seconds before the absolute amplitude decreases exponentially.
[0045] Peak detection algorithms can be used to continuously monitor sampled audio signals. A peak detection algorithm can be a conventional algorithm used, for example, to identify (absolute) peaks in a signal by comparing the amplitude of the sampled audio signal to a predefined threshold. The threshold can represent a normal or expected sound level. When the amplitude of the sampled audio signal exceeds the threshold, an absolute amplitude peak is identified as a peak. The peak detection algorithm can record the occurrence of peaks, including their amplitude and the time of their occurrence. This ensures accurate detection and recording of (absolute) amplitude peaks. Peaks can be used as a criterion for detecting events in the sampled audio signal, which can therefore include two distinct behaviors (i.e., amplitude peaks, followed by exponential behavior).
[0046] The absolute amplitude peak can be at least an order of magnitude higher than the absolute amplitude of the sampled audio signal caused by background noise from the surrounding environment. However, the absolute amplitude peak of an event can, for example, be at least 20 times or two orders of magnitude higher than the absolute amplitude in a sampled audio signal produced by normal or conventional sound. Normal or conventional sound can refer here to an audio signal that follows standard acoustic characteristics or a (statistically) typical amplitude range.
[0047] Therefore, it can provide convenient detection of events, for example, because the absolute amplitude peak of an event can be significantly higher (and therefore more distinguishable or different) compared to other absolute amplitudes of the sampled audio signal.
[0048] The time period can begin at least one second after the absolute amplitude peak is detected.
[0049] In other words, the method could include waiting at least one second after detecting the absolute amplitude peak and / or before performing the detection event step. Therefore, certain audio signal behaviors can be excluded from the analysis.
[0050] In the example, the method could include: waiting up to three seconds after detecting the absolute amplitude peak, and then detecting the event. In other words, the time period can begin up to three seconds after the detection of the absolute amplitude peak.
[0051] The instantaneous rate of change of an exponentially decreasing absolute amplitude can be linearly proportional to the corresponding instantaneous amplitude of the sampled audio signal.
[0052] In other words, during the period of exponential decrease in absolute amplitude, the slope of the amplitude-time curve representing the sampled audio signal can be linearly proportional to the amplitude of the sampled audio signal. This holds true for every time point during the period of exponential decrease.
[0053] A linear ratio between slope and amplitude can be a characteristic of an event. Specifically, a linear ratio between slope and amplitude can be a characteristic of an exponentially decreasing absolute amplitude.
[0054] Therefore, it can provide more accurate detection of events within the sampled audio signal, enabling faster and more accurate detection of microphone device tampering.
[0055] The steps for detecting an event may include: segmenting the sampled audio signal into multiple time slots; and within each time slot: fitting a linear model (i.e., applying linear regression) to capture local trends, determining the average amplitude value of the linear model, determining the slope of the linear model, and determining a fit quality metric of the linear model, such that each time slot is represented by a matrix including the average amplitude value, the slope, and the fit quality metric.
[0056] In other words, the method may include: a smoothed sampled audio signal.
[0057] By representing each time slot with the matrix described above, the sampled audio signal can be represented as a series of values for average amplitude, slope, and fit quality metric.
[0058] In a sense, the sampled audio signal can be divided into time slots, where the sampled audio signal in each time slot is modeled using a linear regression method (where the sampled audio signal in each time slot is estimated to fit a linear equation). For the linear fit in each time slot, the average amplitude value, slope, and quality metric of the fit can be determined.
[0059] Linear models can be based on, for example, the Theil-Sen-Kendall-Sigel method, least squares regression, minimum absolute deviation regression, ridge regression, robust regression, or lasso regression.
[0060] A fit quality metric can be used as an indicator of how linear the sampled audio signal within a time slot is. A fit quality metric can also be called a goodness-of-fit metric.
[0061] The fit quality metric can be, for example, the value of R², adjusted R², RMSE, MAE, AIC, BIC, or any other suitable metric used to express the goodness or quality of the fit. The fit quality metric can be, for example, between 0 and 1, where 1 can represent a perfect fit.
[0062] The method may include: within each time slot, first fitting a linear model, and then using a statistical measure such as the sum of squared residuals to evaluate the deviation of each data point (i.e., magnitude) from the fitted linear model. Data points that deviate from the local trend, such as exceeding a predetermined threshold, can be considered outliers. Outliers can then be excluded from the regression analysis within each time slot.
[0063] Therefore, smoothing can be performed using local linear smoothing with block-truncated least squares regression.
[0064] This truncation process ensures that linear regression is not overly affected by extreme values, resulting in a more reliable and accurate representation of the underlying local trends in the sampled audio signal.
[0065] Fitting a linear model can be, for example, an adaptively truncated least squares model. This can include adjusting the fitting process to minimize the impact of outliers and provide a more robust regression.
[0066] The event detection step may further include: for a continuous time slot sequence, determining whether the fitted quality metric is higher than a threshold for at least a portion of the time slots in the continuous time slot sequence.
[0067] In other words, it can be determined whether the sampled audio signal (i.e., amplitude) approximates a linear line within each time slot, at least a portion of the time slots. Therefore, if the fit quality metric is not higher than a threshold for at least a portion of the time slots in a continuous time slot sequence, then an undetected event can be determined.
[0068] In the example, the threshold can be set to at least 0.7 in the range of 0 to 1, where 1 represents a perfect fit.
[0069] A portion of a time slot may be, for example, at least 20% or at least 10% to 30% of the time slots in a continuous time slot sequence.
[0070] A continuous time-slot sequence can correspond to, for example, 5 to 15 seconds or about 10 seconds.
[0071] For example, if the sampled audio signal approximates a linear line for at least 20% of the time slots within a 10-second interval (i.e., has a sufficiently high measure of fit quality), it can indicate that an event has occurred.
[0072] Generally, adding at least a subset of the fit quality metric that exceeds a threshold helps identify the presence of local linear trends. By setting a threshold, it is beneficial to filter out noise and irrelevant parts or segments from the sampled audio signal. This can lead to more robust detection of events in the sampled audio signal.
[0073] The steps of detecting an event may further include: identifying a time slot interval with the highest sum of values for the best-fit quality metric across multiple time slots; and for each time slot interval, determining a fraction representing the probability that the event has occurred in the sampled audio signal within that time slot interval.
[0074] The time slot range corresponding to the highest (i.e., maximum) value of the fit quality metric can represent most time slots, such as around 70% to 80% or about 75%.
[0075] Time slot intervals can be determined from a continuous sequence of time slots. A time slot interval can be determined, for example, by selecting the top 70% of time slots with the highest fit quality metric within a continuous sequence of time slots. For example, time slots (or matrices corresponding to time slots) can be sorted according to their fit quality metric, such that the bottom 30% with the lowest fit quality metric (i.e., the worst fit) can be excluded from further processing.
[0076] This allows for the exclusion of irrelevant or unlikely-to-correspond-to-event data, resulting in more reliable event detection.
[0077] The score can be determined based on multiple sub-scores, each sub-score corresponding to a specific time slot sub-interval within the time slot interval.
[0078] In other words, a time slot interval can be divided into multiple sub-intervals, allowing a score to be determined for each sub-interval. Sub-intervals can be non-overlapping and / or sequentially shifted within the interval.
[0079] The size of the sub-interval can be chosen arbitrarily and / or fixedly. For example, a time slot interval can be divided into two or more sub-intervals.
[0080] Therefore, the overall (i.e., final) score can be determined based on the scores from the sub-intervals (i.e., sub-scores), for example, by the median or mean.
[0081] This allows for finer-grained analysis of the sampled audio signal, capturing variations within smaller segments that might be overlooked in a broader analysis. By evaluating sub-scores for each sub-interval, the method can more effectively detect local patterns and anomalies (i.e., events).
[0082] Therefore, a more reliable overall score can be provided to indicate the occurrence of an event.
[0083] The event detection step may further include: determining a linear regression that uses the slope of the linear model for each time slot in the time slot interval as a function of the average amplitude value of the linear model for each time slot in the time slot interval, such that k = c1 + m * c0, where k is the slope of the linear model in each time slot, m is the average amplitude value of the linear model in each time slot, and c1 and c0 are constants.
[0084] Therefore, based on the values of slope and average amplitude in the time slot (i.e., in the matrix corresponding to the time slot), a linear relationship between these values can be determined.
[0085] A linear relationship between the average amplitude and slope over a time slot can be a characteristic of an event. Therefore, by determining the existence of such a linear relationship, it can be determined that an event has occurred in the sampled audio signal.
[0086] Linear regression can be called final linear regression or km linear regression, for example, to distinguish it from linear models fitted within each time slot. However, linear regression can benefit from the same discussion as linear models fitted within each time slot.
[0087] Linear regression can be based on, for example, the Theil-Sen-Kendall-Sigel method, least squares regression, minimum absolute deviation regression, ridge regression, robust regression, or lasso regression.
[0088] In the example, linear regression can be determined as a function of the slope of the linear model for each time slot (e.g., for all time slots or a continuous sequence of time slots) as a function of the average magnitude value of the linear model for each time slot (similarly, for all time slots or a continuous sequence of time slots).
[0089] The score can be represented by a quality metric of the fit of a given linear regression k = c1 + m * c0.
[0090] The quality metric for a linear regression fit can be called the final fit quality metric or the km-fit quality metric, for example, to distinguish the km-fit quality metric from the fit quality metric for a linear model within a time slot. However, the km-fit quality metric can benefit from the same discussion as the fit quality metric for a linear model within a time slot.
[0091] The score can be compared to a fixed threshold. The fixed threshold can be set to at least 0.7, for example, in the range of 0 to 1, where 1 represents a perfect fit.
[0092] In the example, a threshold of 0.7 can result in no false triggers for more than 20 billion 10-second analog signals.
[0093] Therefore, when the score is above a fixed threshold (i.e., when the km linear regression results in a sufficiently high quality metric for the km fit), it can be determined that the event has occurred in the sampled audio signal.
[0094] Fractions can be determined using fractional functions. In the example, when the time interval is divided into multiple sub-intervals, the fraction can be determined, for example, as:
[0095] Score = F(Q of SLR) SLR ([k=c*m] over all time slots in <time slot sub-interval>)),
[0096] Wherein, SLR is a simple linear regression (i.e., km linear regression) of the slope k and average amplitude m for each time slot in the time slot sub-interval, Q SLR F is a quality measure of the SLR fit for each subinterval, and F is a function of the mean or class median.
[0097] In other words, the score can be represented by the value of a quality metric (i.e., sub-score) of the fit of the km linear regression for the time slot sub-intervals (e.g., median, mean, etc.). In a sense, the score can be given by score = F(sub-score), that is, the median or mean of the sub-scores for each sub-interval.
[0098] The score can be defined, for example, as the second largest value of the quality metric of the km fit within the subinterval.
[0099] If no fraction or sub-fraction is found, the fraction can be set to 0.
[0100] The computer-implemented method may further include performing verification steps, including determining whether c0 corresponds to the zero state of the sampled audio signal, and determining whether c1 corresponds to the expected time period during which the absolute magnitude of the event decreases exponentially.
[0101] Therefore, it can be determined that the event did not occur when c0 does not correspond to the zero state of the sampled audio signal (i.e., the DC of the microphone) and / or when c1 does not correspond to the expected time period during which the magnitude of the event decreases exponentially.
[0102] The constant c0 may, for example, preferably match or correspond to a zero state over at least the past 60 seconds (i.e., the sampled audio signal DC).
[0103] The expected time period can be a known or predefined time period. In particular, the expected time period can be a characteristic feature of different types of microphones or different types of microphone devices. The expected time period can be determined, for example, through experimentation or a lookup table.
[0104] According to the second aspect, a non-transitory computer-readable medium is provided, on which instructions are stored, which, when executed by a processor, cause the processor to perform the steps of the method according to the first aspect.
[0105] This aspect generally presents the same or corresponding advantages as the first aspect.
[0106] According to a third aspect, a system is provided, comprising: a microphone device including a microphone disposed in a cavity that opens to the surrounding environment of the microphone device through a microphone hole; and a processing unit configured to perform the method according to the first aspect.
[0107] This aspect generally presents the same or corresponding advantages as the first aspect.
[0108] The processing unit can be configured to receive audio signals from the microphone device and perform various processing tasks as well as extract relevant information. Upon receiving an analog audio signal, the processing unit can digitize the analog audio signal.
[0109] The processing unit can analyze the audio signal to detect or separate specific audio features such as events. The processing unit can employ filtering techniques to remove unwanted frequencies and enhance desired audio components such as low-frequency components.
[0110] The processing unit can be, for example, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA).
[0111] Furthermore, the processing unit can communicatively connect to multiple microphone devices, such as a network of microphone devices. Therefore, the processing unit can simultaneously detect or monitor tampering with multiple microphone devices.
[0112] The processing unit can be located away from the microphone assembly. However, it should be recognized that the processing unit can be integrated into the microphone assembly.
[0113] Generally, unless otherwise expressly defined herein, all terms used in the claims shall be interpreted according to their ordinary meaning in the art. Unless otherwise expressly stated, all references to “a” / “an” / “the” [element, device, component, apparatus, step, etc.] shall be openly interpreted as referring to at least one instance of said element, device, component, apparatus, step, etc. Unless expressly stated otherwise, the steps of any method disclosed herein need not be performed in the exact order disclosed. Attached Figure Description
[0114] The above and other objects, features, and advantages of the invention will be better understood from the following illustrative and non-limiting detailed description of preferred embodiments of the invention with reference to the accompanying drawings, wherein like reference numerals will be used for similar elements, wherein:
[0115] Figure 1 A diagram illustrating a computer-implemented method for detecting tampering with a microphone device.
[0116] Figures 2A to 2B The schematic diagram shows a microphone device including a microphone in a cavity.
[0117] Figure 3 The schematic diagram shows a system including a microphone assembly and a processing unit.
[0118] Figure 4A The illustration includes an exemplary audio signal corresponding to an event of tampering with the microphone device.
[0119] Figure 4B The illustration includes another exemplary audio signal corresponding to an event of tampering with the microphone device.
[0120] Figure 4C The illustration includes exemplary audio signals corresponding to multiple events of tampering with the microphone device. Detailed Implementation
[0121] The invention will now be described more fully with reference to the accompanying drawings, in which presently preferred embodiments of the invention are illustrated. However, the invention may be embodied in many different forms and should not be construed as limited to the embodiments set forth herein; rather, these embodiments are provided for thoroughness and completeness and to fully convey the scope of the invention to those skilled in the art.
[0122] Figure 1 A block diagram illustrating a computer-implemented method 1000 for tamper detection of a microphone device 100. Method 1000 includes:
[0123] Receive audio signal 112 from microphone 110;
[0124] The audio signal 112 is sampled for the low-frequency band by 1200.
[0125] Detect event 1300 140, which includes an exponentially decreasing absolute magnitude of the audio signal 130 sampled over a period of at least one second; and
[0126] When event 140 is detected, output 1400 tamper detection signal.
[0127] Despite Figure 1 As not shown above, but the steps of detecting event 140 of 1300 may include: segmenting the sampled audio signal 130 into multiple time slots; and within each time slot: fitting a linear model to capture local trends, determining the average amplitude value of the linear model, determining the slope of the linear model, and determining a fit quality metric of the linear model, such that each time slot is represented by a matrix including the average amplitude value, the slope, and the fit quality metric.
[0128] Furthermore, the steps for detecting 1300 events may include: for a continuous time slot sequence, determining whether the fitted quality metric is higher than a threshold for at least a portion of the time slots in the continuous time slot sequence.
[0129] The step of detecting event 140 may further include: identifying a time slot interval with the highest sum of values for the best-fit quality metric among multiple time slots; and for each time slot interval, determining a fraction representing the probability that event 140 has occurred in the sampled audio signal 130 within the time slot interval.
[0130] exist Figure 1 In another example not depicted, the step of detecting event 1300 140 may include: determining a linear regression that is a function of the slope of the linear model for each time slot in the time slot interval as a function of the average amplitude value of the linear model for each time slot in the time slot interval, such that k = c1 + m * c0, where k is the slope of the linear model for each time slot, m is the average amplitude value of the linear model for each time slot, and c1 and c0 are constants.
[0131] It should be recognized that a score can be determined based on multiple sub-scores, each corresponding to a specific time slot sub-interval within a time slot interval. In a particular example, the score can be represented by a quality metric of the fit to a defined linear regression k = c1 + m * c0.
[0132] Furthermore, despite Figure 1 Not shown, but method 1000 may further include: performing a verification step, including: determining whether c0 corresponds to the zero state 132 of the sampled audio signal 130, and determining whether c1 corresponds to the expected time period during which the absolute magnitude of the event decreases exponentially.
[0133] Figure 2A The figure shows a partial cross-section of a microphone device 100. The microphone device 100 includes a microphone 110 disposed in a cavity 120 of the microphone device 100. The cavity 120 is open to the surrounding environment of the microphone device 110 through a microphone hole 122.
[0134] Note that microphone hole 122 is the only opening in cavity 120. Therefore, microphone 110 is placed in cavity 120, which is acoustically isolated from the rest of microphone device 100. Microphone 110 can thus effectively capture sound from its surrounding environment while mitigating interference from internal components.
[0135] Generally speaking, placing the microphone 110 inside the cavity 120 helps protect the microphone 110 from external environmental factors such as dust and moisture.
[0136] Furthermore, despite Figure 2A Although not shown, it should be understood that microphone 110 may be electrically connected to internal components of microphone device 120.
[0137] Furthermore, it should be recognized that the microphone hole 122 (i.e., the opening of the cavity 120) may include a mesh or protective mesh or layer. This can reduce the entry of dust and debris into the cavity 120, thereby protecting the microphone 110 and still ensuring clear audio capture.
[0138] Figure 2BThe diagram shows a microphone device 100 in the form of a camera device. (Regarding...) Figure 2A Similar to the description, microphone 110 is integrated into the camera unit. (As such...) Figure 2B As can be seen, microphone 110 is arranged in cavity 120, which is open to the surrounding environment through microphone hole 122.
[0139] In a sense, the microphone 110 is embedded in a dedicated cavity 120 within the camera housing of the camera device, and has an opening 122 to the external environment.
[0140] although Figures 2A to 2B Not shown, but microphone device 100 may include a processor configured to execute computer-implemented methods for detecting tampering with microphone device 110 (i.e., blockage of microphone hole 122), for example, as per [reference to...]. Figure 1 As described.
[0141] It should be further understood that the proportions, shapes, and relative scales in the accompanying drawings are exemplary and have been enlarged to aid visualization. In practice, the microphone hole 122 may have a relatively small diameter, for example, about 1 mm. Similarly, the depth of the microphone hole may be in the range of about 1 mm to 2 mm.
[0142] Figure 3 The illustrated system 300 includes: a microphone device 100 comprising a microphone 110 disposed in a cavity 120 open to the surrounding environment of the microphone device 110 through a microphone hole 122; and a processing unit 200 configured to perform a computer-implemented method for detecting tampering with the microphone device 110, such as, as per [reference to...] Figure 1 As described.
[0143] The microphone device 100 is illustrated here as a cross-section of the microphone device 100. The microphone device 100 may be, for example, about... Figures 2A to 2B Any of the microphone devices discussed.
[0144] Further in Figure 3 In the cavity 120, the microphone hole 122 is covered by an object 400. The object 400 placed on the microphone hole 122 of the cavity 120 can be similar to the microphone device 100 that is being tampered with.
[0145] The object 400 can be any suitable object covering the microphone hole 122, such as chewing gum, cloth, tape, or a hand. The object 400 arranged above the microphone hole 122 can generally correspond to the action of the microphone 110 designed to block sound from reaching the microphone device 100.
[0146] The object 400 can press against the microphone hole 122, thereby isolating the cavity 120 from the surrounding environment; that is, the object 400 can provide a tight seal for the cavity 120.
[0147] exist Figure 3 In the process, the processing unit 200 receives audio signal 112 from the microphone device 100.
[0148] At processing unit 200, audio signal 112 can be received as an analog audio signal. Therefore, processing unit 200 can be configured to digitize the analog audio signal.
[0149] The processing unit 200 can analyze the audio signal 112 to detect tampering with the microphone device 100 (i.e., detect events of an exponentially decreasing absolute magnitude of the audio signal 130 sampled during a time period).
[0150] Processing unit 200 may employ filtering techniques to remove unwanted frequencies and enhance desired audio components such as low-frequency components. In particular, processing unit 200 may include a low-frequency band. The low-frequency band may include frequencies below 500 Hz. However, it should be understood that the low-frequency band may be part of microphone device 100 or microphone 110.
[0151] The processing unit 200 can be, for example, a microprocessor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or a field-programmable gate array (FPGA). Generally, the processing unit 200 can be any suitable device with processing capabilities.
[0152] Furthermore, despite Figure 3 While not explicitly described, processing unit 200 can communicatively connect to multiple microphone devices 100, such as a network of microphone devices 100. In other words, processing unit 200 can be configured to simultaneously analyze audio signals 130 from multiple microphone devices 100. Processing unit 200 (e.g., a single processing unit 200) can thus detect or monitor tampering with multiple microphone devices.
[0153] In addition, Figure 3 In this design, the processing unit 200 is depicted as being located remotely from the microphone device 100. However, the processing unit may, for example, be integrated into the microphone device 100. In this example, the main microphone device in a network of microphone devices may include the processing unit 200 or have processing capabilities. Therefore, audio signals from other microphone devices in the network of microphone devices can be sent to the main microphone device for processing.
[0154] System 300 may be, for example, an audio monitoring system or part of an audio monitoring system. The system may further include means for presenting an alarm or indication to the user or supervisor of the audio monitoring system when the microphone device is tampered with (e.g., via a monitor).
[0155] although Figure 3 Not explicitly shown, but a non-transitory computer-readable medium may be provided storing instructions that, when executed by a processor (such as processing unit 200), cause the processor to perform steps of a computer-implemented method for tamper detection of the microphone device 100, for example, as per [reference to...]. Figure 1 As described.
[0156] Figure 4A An exemplary sampled audio signal 130 from a tampered microphone device is shown. The sampled audio signal 130 is represented here as a graph of amplitude changing over time.
[0157] The first part of the sampled audio signal 130 corresponds to the general audio characteristics or typical behavior of the audio signal, such as background noise.
[0158] A peak amplitude was observed in the second part of the sampled audio signal 130.
[0159] The peak amplitude can represent the maximum absolute amplitude value reached by the sampled audio signal 130 within a portion of the sampled audio signal 130. In a sense, the peak amplitude can be regarded as the maximum absolute amplitude value.
[0160] However, in Figure 4A In this context, the peak amplitude also includes negative values. Therefore, the peak amplitude can be called the absolute peak amplitude. Consequently, the trough or minimum amplitude value can also be included in the absolute peak amplitude.
[0161] The (absolute) peak amplitude is significantly higher than the amplitude at other locations on the sampled audio signal 130. Typically, the peak amplitude can be at least an order of magnitude higher than the amplitude of the sampled audio signal 130 caused by background noise from the surrounding environment.
[0162] exist Figure 4A In the specific audio signal 130 described herein, after the absolute amplitude peak (specifically after the amplitude minimum), the sampled audio signal 130 increases toward and exceeds the zero state 132. Subsequently, the sampled audio signal 130 reaches a local maximum.
[0163] Following this local maximum, the amplitude of the sampled audio signal 130 was observed to decrease exponentially. The amplitude of the sampled audio signal 130 decreased exponentially during the time period P.
[0164] The magnitude of the exponential decrease during its period can typically be between 1 and 20 seconds.
[0165] exist Figure 4A In this context, time period P begins approximately three seconds after the peak of the absolute magnitude. However, time period P (i.e., exhibiting an exponentially decreasing absolute magnitude) can begin directly (e.g., ...). Figure 4B (as described in the text) or begins at least one second after the peak of the absolute amplitude.
[0166] A peak amplitude and / or an exponentially decreasing amplitude (i.e., a ramp) may be referred to as event 140. Event 140 may indicate tampering with microphone device 100.
[0167] exist Figure 4A The upper left illustration schematically depicts an exemplary alteration of the microphone device 100. Here, an object 400 presses against the microphone hole 122 of the microphone device to block sound from reaching the microphone 110 of the microphone device 100.
[0168] like Figure 4A As illustrated, the act of blocking the microphone hole 122 or the occurrence of an event 140 in the sampled audio signal 130 is captured by the microphone 110.
[0169] The slight movement or vibration of the microphone device 100 caused by pressing the object 400 against the microphone hole 122 (e.g., a slight tremor of the hand during the movement) affects the amplitude of the audio recorded by the microphone 110. This movement or vibration can cause an absolute amplitude peak to appear in the sampled audio signal 130. In particular, in the case of tampering, a higher frequency or amplitude value can be detected from the touch of the microphone device 100.
[0170] Therefore, the amplitude peaks seen here comprise multiple amplitude peaks. The low resolution of the sampled audio signal at the amplitude peaks is caused by the strong and rapid movement of the microphone diaphragm during the action of attaching the object 400 to the microphone device 100 (e.g., due to vibration of the microphone device 100). The amplitude peaks can therefore be referred to as root mean square (RMS) peaks (i.e., the peak of many peaks).
[0171] In this context, the term "RMS peak" can be used to describe the effective amplitude of a sampled audio signal 130 exhibiting multiple peaks that form a larger peak. The RMS value can be represented by the measurement of the amplitude peak of the sampled audio signal 130 by calculating the square root of the square mean of the instantaneous values over a specified time period (i.e., over the many amplitude peaks caused by the touch microphone device 100).
[0172] Furthermore, by covering the microphone hole 122, air is trapped in the cavity and cannot escape, or escapes more slowly than before the microphone hole was covered. These physical interactions with the microphone device 100 (i.e., tampering) can cause event 140 in the sampled audio signal 130 recorded by the microphone device 100.
[0173] Event 140 can be characterized by an instantaneous rate of change that decreases exponentially and is linearly proportional to the corresponding instantaneous amplitude of the sampled audio signal 130.
[0174] exist Figure 4A The upper right inset depicts the amplified portion of the sampled audio signal 130. Specifically, it can be seen that the sampled audio signal 130 in time slot T is approximated by a linear model. The linear model for each time slot, and thus the sampled audio signal 130, is represented here by matrix M. Matrix M includes the values corresponding to the slope k, the average amplitude value M, and the fit quality metric Q of the linear model.
[0175] In the example, each time slot T in the interval I of n time slots n It can be derived from the corresponding matrix M n =(k n ,m n Q n )express.
[0176] The fit quality metric Q of time slot T can be used to identify intervals I or time periods P in which the sampled audio signal 130 decreases exponentially during its duration. Specifically, if the fit quality metric Q indicates a poor fit for most (or, for example, more than 20%) of the consecutive time slots T (e.g., Q < 0.7 in the range of 0 to 1), it can be determined that event 140 has not occurred. On the other hand, if most (e.g., at least 80%) of the consecutive time slots have a fit quality metric Q indicating an accurate fit (e.g., Q > 0.7), the sampled audio signal 130 can be further processed to determine whether the event has occurred.
[0177] although Figure 4A It is not explicitly described, but time slot interval I can constitute multiple sub-intervals of time slot T.
[0178] The average amplitude value m and slope value k of each time slot in interval I or subinterval can be used to form a (km) linear regression (i.e., k = c1 + m * c0).
[0179] The quality or goodness of linear regression (e.g., Rm) of matrix data (km) from time slot T in interval I or subinterval I. 2The value can correspond to a score. A score can represent the probability that event 140 has occurred. When determining a score for each sub-interval of time slot T, the mean or median function can be used to represent the resulting score values for multiple scores. For example, if R is determined for each (mk) linear regression of each sub-interval... 2 The value can then be determined based on R. 2 The values are calculated as the average or median. However, for example, the largest or second largest R... 2 Values can be used to represent fractions.
[0180] When considering R for quality metrics 2 When the value is set, it has been determined that for simulated data, the score is typically greater than 0.999. It was observed that this decreases to greater than 0.95 with smaller model biases, and further decreases to greater than 0.7 with added noise.
[0181] Therefore, for scores above 0.7 (in the range of 0 to 1, where 1 represents the full score), it can be determined that event 140 has occurred (i.e., the microphone device 100 has been tampered with).
[0182] In particular, in simulations with added noise, it has been observed that a fractional threshold of 0.7 can result in no erroneous triggering of audio signals after more than 20 billion 10-second simulated samples.
[0183] exist Figure 4A The sampled audio signal 130 and event 140 have been described when event 140 forms a positive amplitude value (i.e., on the positive side of the zero state 132 (i.e., DC) of the sampled audio signal). However, it should be understood that event 140 can also form a negative amplitude value (i.e., on the negative side of the zero state 132). Therefore, the amplitude peak and the exponential decrease in amplitude can generally be referred to as the absolute amplitude peak and the exponential decrease in absolute amplitude, respectively.
[0184] Figure 4B The illustration includes another sampled audio signal 130 corresponding to an event of tampering with the microphone device.
[0185] Figure 4B To a large extent benefited Figure 4A The discussion. However, in Figure 4B In the middle, the amplitude peak is immediately followed by an exponentially decreasing amplitude. Specifically, the amplitude decreases exponentially toward the zero state 132 of the sampled audio signal 130.
[0186] In the example, the event could include the peak absolute amplitude of the sampled audio signal 130, followed directly by the exponentially decreasing absolute amplitude of the sampled audio signal during the time period P.
[0187] Therefore, the detection of tampering with the microphone device can correspond to the detection of the absolute amplitude peak, followed by the exponentially decreasing absolute amplitude of the audio signal 130 sampled during the time period P. The exponentially decreasing behavior of the sampled audio signal 30 (i.e., the time period P) can thus begin directly after the amplitude peak.
[0188] Figure 4C This shows that the sampled audio signal 130 from the microphone device has been tampered with at multiple points in time. It is recognized that... Figure 4C Benefiting Figure 4A and Figure 4B The discussion.
[0189] exist Figure 4C In this context, the characteristic behavior describing an exponentially decreasing absolute amplitude value can occur on either side of the zero state 132. Specifically, the event can take various forms; however, each event corresponding to the tampering includes an exponentially decreasing absolute amplitude of the audio signal 130 sampled during the time period P.
[0190] It will be appreciated that the invention is not limited to the embodiments shown. Therefore, various modifications and variations are conceived within the scope of the invention as defined by the appended claims.
Claims
1. A computer-implemented method (1000) for tamper detection of a microphone device (100), the microphone device (100) comprising a microphone (110) disposed in a cavity (120) open to the surrounding environment of the microphone device (110) through a microphone hole (122), the method (1000) comprising: Receive (1100) audio signal (112) from the microphone (110); The audio signal (122) is sampled (1200) for the low frequency band by decimating the audio signal (112) to reduce the sampling rate. The detection (1300) includes an event (140) of an exponentially decreasing absolute amplitude of the sampled audio signal (130), wherein the exponentially decreasing absolute amplitude of the event (140) decreases exponentially over a time period (P) of at least one second; and When the event (140) is detected, a tamper detection signal (1400) is output.
2. The computer-implemented method (1000) according to claim 1, wherein, The low-frequency band includes frequencies below 500 Hz.
3. The computer-implemented method (1000) according to claim 1, wherein, The absolute magnitude of the event (140) decreasing exponentially during the time period (P) during which the magnitude decreases exponentially is 1 to 20 seconds.
4. The computer-implemented method (1000) according to claim 1, further comprising: Before detecting (1300) the event (140) of the sampled audio signal (130) with the exponentially decreasing absolute amplitude, the peak absolute amplitude of the sampled audio signal (130) is detected, wherein the peak absolute amplitude is at least one order of magnitude higher than the absolute amplitude of the sampled audio signal (130) caused by the background noise of the surrounding environment.
5. The computer-implemented method (1000) according to claim 4, wherein, The time period (P) begins at least one second after the absolute amplitude peak is detected.
6. The computer-implemented method (1000) according to claim 1, wherein, The instantaneous rate of change of the absolute amplitude, which decreases exponentially, is linearly proportional to the instantaneous amplitude of the sampled audio signal (130).
7. The computer-implemented method (1000) according to claim 1, wherein, The steps for detecting (1300) the event (140) include: The sampled audio signal (130) is segmented into multiple time slots; and Within each time slot: Fitting a linear model to capture local trends Determine the average magnitude value of the linear model. Determine the slope of the linear model, and Determine the fit quality metric for the linear model such that each time slot is represented by a matrix including the average amplitude value, slope, and fit quality metric.
8. The computer-implemented method (1000) according to claim 7, wherein, The step of detecting (1300) the event further includes: For a continuous time-slot sequence, determine whether the fit quality metric is higher than a threshold for at least a portion of the time slots in the continuous time-slot sequence.
9. The computer-implemented method (1000) according to claim 7, wherein, The step of detecting (1300) the event (140) further includes: Identify the time slot interval with the highest sum of values for the best fit quality metric among the plurality of time slots; and For the time slot interval, a fraction representing the probability that the event (140) has occurred in the sampled audio signal (130) within the time slot interval is determined.
10. The computer-implemented method (1000) according to claim 9, wherein, The score is determined based on multiple sub-scores, each sub-score corresponding to a specific time slot sub-interval within the time slot interval.
11. The computer-implemented method (1000) according to claim 9, wherein, The step of detecting (1300) the event (140) further includes: A linear regression is determined that uses the slope of the linear model for each time slot in the time slot interval as a function of the average amplitude value of the linear model for each time slot in the time slot interval, such that k = c1 + m * c0, where k is the slope of the linear model for each time slot, m is the average amplitude value of the linear model for each time slot, and c1 and c0 are constants.
12. The computer-implemented method (1000) according to claim 11, wherein, The score is represented by a quality metric for the fit of the determined linear regression k=c1+m*c0.
13. The computer-implemented method (1000) according to claim 11, further comprising: Perform the verification steps, including: Determine whether c0 corresponds to the zero state (132) of the sampled audio signal (130), and Determine whether c1 corresponds to the expected time period during which the absolute magnitude of the event decreases exponentially.
14. A non-transitory computer-readable medium storing instructions that, when executed by a processor, cause the processor to perform the steps of the method (1000) according to any one of claims 1 to 13.
15. A system (300) comprising: A microphone device (100) includes a microphone (110) disposed in a cavity (120) that opens to the surrounding environment of the microphone device (110) through a microphone hole (122); and The processing unit (200) is configured to perform the method (1000) according to any one of claims 1 to 13.