System and method for property protection based on audio and acceleration related data
By using audio data and acceleration-related data of the vehicle subject and combining machine learning models to classify and verify vehicle tampering events, it solves the problem that it is difficult to detect vehicle tampering events without relying on clear vision in the prior art, and achieves rapid and accurate detection of specific types of vehicle tampering events.
Patent Information
- Application Number
- CN202411893831.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-12-21
- Filing Date
- 2024-12-20
- Publication Date
- 2025-06-24
AI Technical Summary
The prior art is difficult to effectively detect and prevent vehicle tampering events without relying on clear line of sight, especially thefts occurring in the loading space of a van or under the vehicle's underside.
By combining audio data with acceleration-related data of vehicle subjects, using machine learning models (such as the time convolutional network model) for classification, potential vehicle tampering events are identified, and false positives are reduced through acceleration-related data verification.
It realizes rapid and effective detection of certain types of vehicle tampering events, reduces false alarms caused by other noise events, and improves the accuracy and reliability of the system.
Smart Images

Figure CN120197228A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to systems and methods for property protection based on audio and acceleration-related data. Background Art
[0002] An accelerometer is a device that measures acceleration-related data of a body to which the accelerometer is mounted. The acceleration-related data may include the relative acceleration of the body (i.e., the relative rate of change of velocity) and / or the relative jerk of the body (i.e., the relative rate of change of acceleration).
[0003] A machine learning model may refer to an algorithm-based computer program that is trained to identify patterns in data and make predictions or classifications based on such learned pattern identification. Summary of the Invention
[0004] Thieves and other criminals often target vehicles with the aim of entering and stealing goods from inside the vehicle. Thieves use various techniques (such as using manual tools (e.g., drills), lock picking, keyless entry, etc.) to enter the vehicle. A recent technique (referred to as "peel and steal") involves prying / pulling the top of the vehicle door outward and downward, or otherwise peeling the vehicle body apart.
[0005] Other types of vehicle tampering events involve thieves stealing valuable vehicle components. For example, catalytic converter theft has become a serious problem. According to some statistics, there were over 14,400 catalytic converter thefts in 2020, an increase of over 1000% compared to 2018. Since the cost of replacing a catalytic converter per vehicle can be as high as $6,000, this increase in catalytic converter theft is a major concern for many businesses. This concern is particularly acute for businesses that have fleets of vehicles parked adjacent to each other. Thieves often target these fleets and steal catalytic converters from multiple vehicles during related theft incidents. Van fleets are particularly vulnerable to this type of theft due to their relatively high vehicle floors, which allow thieves to reach the catalytic converters that are typically installed on the underside of the vehicle more quickly / easily.
[0006] Various existing technologies have been deployed to detect and prevent the above-mentioned vehicle tampering events (as used herein, vehicle tampering events may refer to the occurrence of unauthorized physical tampering / interference with a vehicle). Examples of these technologies include the use of cameras (e.g., dashboard cameras of vehicles), computer vision, proximity sensors (e.g., radar, lidar, and / or sonar sensors of vehicles), and passive infrared sensors, etc. However, these existing technologies may not be suitable for detecting / preventing certain types of vehicle tampering events because they typically rely on a clear line of sight to the thief / vehicle tampering event.
[0007] For example, when a burglar enters the loading space of a van, the line of sight to the burglar / vehicle tampering event is typically blocked or obstructed by the metal walls of the van. Similarly, when a burglar lies beneath the underside of a vehicle to cut off / saw off the catalytic converter, the line of sight to the burglar / theft event is also typically blocked or obstructed. Thus, prior art that relies on a clear line of sight to detect vehicle tampering events (e.g., cameras, computer vision, proximity sensors, passive infrared sensors, etc.) generally cannot detect these types of vehicle tampering events, or, relatedly, cannot detect these types of vehicle tampering events quickly enough to achieve optimal deterrence.
[0008] In this context, examples of the presently disclosed technology provide innovative systems and methods for detecting vehicle tampering events without relying on a clear line of sight to the vehicle tampering event. That is, the examples leverage the insight that many types of vehicle tampering events have unique audio signatures. Thus, the examples detect / classify vehicle tampering events based on these unique audio signatures. In some embodiments, the examples may train and deploy a machine learning model (sometimes referred to herein as an “audio model”) to perform such audio-based classification. Additionally, the examples may analyze acceleration-related data (e.g., relative acceleration data of the body of the vehicle, relative jerk data of the body of the vehicle, etc.) to determine suspicious movement of the body of the vehicle (e.g., the walls or doors of the vehicle, the partitions of the vehicle, etc.) during a potential / suspicious vehicle tampering event. Such acceleration-related data may be obtained by accelerometers mounted to the body of the vehicle. This acceleration-related verification step can reduce the occurrence of audio-based false alarm classifications caused by other noise events near the vehicle that have audio signatures similar to vehicle tampering events (e.g., drilling or other noise from a construction site, rain, etc.).
[0009] For example, the alarm system of the presently disclosed technology can operate to: (1) provide audio data from a potential tampering event involving a vehicle to a machine learning model trained with audio signatures of known vehicle tampering events (e.g., handle pull tampering events, drilling tampering events, key-lock tampering events, metal stripping related tampering events, catalytic converter theft tampering events, etc.); (2) in response to the machine learning model classifying the potential tampering event as a vehicle tampering event based on the audio data, compare acceleration-related data (e.g., relative acceleration data or relative jerk data) from a body of the vehicle (e.g., a wall of the vehicle, a partition of the vehicle, etc.) during the potential tampering event with a threshold (e.g., a threshold acceleration value or a threshold jerk value); (3)(a) in response to determining that the acceleration-related data exceeds the threshold, place the alarm system on high alert based on the vehicle tampering event classification (e.g., activate additional sensors of the alarm system, activate an audio alarm, activate a visual alarm, send an alarm notification to a location remote from the alarm system, etc.); and (3)(b) in response to determining that the acceleration-related data does not exceed the threshold, discard the vehicle tampering event classification and maintain the default alert state of the alarm system. Using audio- and vehicle-body-acceleration-based classification that does not rely on a clear line of sight to the burglar / vehicle tampering event, the alarm system can detect certain types of vehicle tampering events (e.g., thefts occurring in the loading space of a van, catalytic converter thefts occurring under the underside of a vehicle, etc.) faster and more effectively than existing / alternative technologies. Additionally, using the above-described vehicle-body-acceleration-related verification step, the alarm system can reduce the occurrence of false alarm classifications caused by other noise events near the vehicle that have audio signatures similar to vehicle tampering events (e.g., drilling or other noise from a construction site, rain, etc.). Reducing the occurrence of false alarm classifications has many advantages, including: (a) increasing consumer trust in the alarm system; (b) reducing the annoyance of false alarms; (c) saving power in embodiments where additional portions of the alarm system are awakened / activated in response to a verified detection / classification of a vehicle tampering event; etc.
[0010] Another intelligent insight exploited by examples of the presently disclosed technology is that certain vehicle tampering events also have unique temporal characteristics. For example, handle pull tampering events (and associated audio from handle pull tampering events) typically have a duration of a few seconds or less. In contrast, catalytic converter theft tampering events (and associated audio from catalytic converter theft tampering events) typically have a duration of 30 seconds to a minute. Based on this insight, examples of the presently disclosed technology can improve the accuracy of audio-based classification by leveraging a temporal convolutional network (TCN) model that is particularly adapted to learn the temporal characteristics as well as the audio characteristics of vehicle tampering events. Relatedly, examples of the presently disclosed technology can convert audio data (e.g., raw audio data or preprocessed audio data) into a temporal format that is more effectively / efficiently processed by the TCN model. By way of example, an example can: (1) receive first audio data (e.g., raw or preprocessed audio data) from a potential vehicle tampering event (e.g., from an audio sensor located within the interior space of a vehicle); (2) encode the first audio data into a latent representation of the first audio data (i.e., a lower-dimensional representation of the first audio data that captures the important / critical features of the first audio data); (3) divide the latent representation into time window frames and stack the time window frames to generate temporal audio data; and (4) provide the temporal audio data to the TCN model for vehicle tampering event classification. By using temporal audio data and the TCN model, examples of the presently disclosed technology can improve classification accuracy relative to alternative methods that lack such a temporal approach. Such an improvement in classification accuracy has similar advantages as described above, including: (a) increasing consumer trust; (b) reducing the annoyance of false alarms; (c) saving power in embodiments where an additional portion of an alarm system is activated / woken in response to a verified detection / classification of a vehicle tampering event; (d) allowing for a customized response to a particular type of vehicle tampering event; and so on.
[0011] Examples of the presently disclosed technology also utilize innovative methods for training a machine learning model (i.e., an audio model) to classify specific types of vehicle tampering events based on audio data. For example, an example can: (1) receive first audio data from known catalytic converter theft tampering events (such known catalytic converter theft tampering events can include “imitated” / “simulated” catalytic converter thefts performed on a vehicle in a controlled scenario and actual catalytic converter thefts “in the wild” detected / classified by an alarm system deployed by the presently disclosed technology); (2) process the first audio data for improved machine learning model training; and (3) use the processed audio data to train a machine learning model (e.g., a TCN model) to classify potential vehicle tampering events as catalytic converter theft tampering events. Processing the first audio data can include various types of processing, including any combination of the following: (a) purifying the first audio data to remove foreign noise or otherwise discard noisy data, (b) labeling the first audio data as related to a catalytic converter theft tampering event, (c) encoding the first audio data into a latent representation, and (d) preparing temporal audio data by dividing the latent representation of the first audio data into time window frames and stacking the time window frames to generate the temporal audio data. Examples can also utilize audio data from known non-catalytic converter theft tampering events (e.g., handle pull tampering events, drilling tampering events, key-lock tampering events, metal stripping related tampering events, etc.) and known non-tampering events (e.g., construction site noise, rain, etc.) to train a machine learning model to classify different types of vehicle tampering events and distinguish them from non-tampering events.
[0012] As described above, examples of the presently disclosed technology provide numerous advantages over existing and potential alternative technologies. For example, by utilizing audio and vehicle body acceleration-based classification that does not rely on a clear line of sight to a burglar / vehicle tampering event, the examples can detect certain types of vehicle tampering events (e.g., thefts occurring within the loading space of a van, catalytic converter thefts occurring beneath the underside of a vehicle, etc.) faster and more efficiently than existing / alternative technologies. Additionally, by utilizing the vehicle body acceleration-related verification steps described above, the examples can reduce the occurrence of false alarm classifications caused by other noise events (e.g., drilling or other noise from a construction site, rain, etc.) near the vehicle that have audio characteristics similar to those of a vehicle tampering event. Reducing the occurrence of false alarm classifications has many advantages, including: (a) increasing consumer trust; (b) reducing the annoyance of false alarms; (c) saving power in embodiments where additional portions of an alarm system are activated / woken in response to a verified detection / classification of a vehicle tampering event; and so on. Correlatively, by using temporal audio data and a TCN model, the examples can improve classification accuracy relative to alternative methods that lack such a temporal approach. Such an improvement in classification accuracy has similar advantages as described above, including: (a) increasing consumer trust; (b) reducing the annoyance of false alarms; (c) saving power in embodiments where additional portions of an alarm system are activated / woken in response to a verified detection / classification of a vehicle tampering event; (d) allowing for customized responses to specific types of vehicle tampering events; and so on.
[0013] While the specific examples detailed herein apply to vehicle protection, it should be understood that the principles disclosed herein can be applied to other types of property. For example, certain types of property may be located (or otherwise stored) within an enclosed space where the line of sight to the property and / or potential tampering events will be blocked or obstructed. For example, the alarm system of the presently disclosed technology can operate to: (1) provide audio data from a potential tampering event to a machine learning model trained with audio characteristics of known tampering events (such as vehicle tampering events or other types of tampering events involving personal property); (2) in response to the machine learning model classifying the potential tampering event as a tampering event based on the audio data, compare acceleration-related data (such as relative acceleration data or relative jerk data) from a subject of the property (such as the subject of the property protected by the alarm system) during the potential tampering event to a threshold (such as a threshold acceleration value or a threshold jerk value); (3)(a) in response to determining that the acceleration-related data exceeds the threshold, place the alarm system on high alert based on the tampering event classification (such as activating additional sensors of the alarm system, activating an audio alarm, activating a visual alarm, sending an alarm notification to a location away from the alarm system, etc.); and (3)(b) in response to determining that the acceleration-related data does not exceed the threshold, discard the tampering event classification and maintain the default alert state of the alarm system. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] The present disclosure is described in detail below with reference to the following drawings. The drawings are provided for illustrative purposes only and depict examples only.
[0015] Figure 1 The architecture of an example alarm system according to various examples of the presently disclosed technology is shown.
[0016] Figure 2 An example method for vehicle tampering event classification and verification according to various examples of the presently disclosed technology is depicted.
[0017] Figure 3 Example operations for performing vehicle tampering event classification and verification according to various examples of the presently disclosed technology are depicted.
[0018] Figure 4 Example operations for classifying a vehicle tampering event based on audio data according to various examples of the presently disclosed technology are depicted.
[0019] Figure 5 Example operations for classifying a tampering event involving property based on audio data according to various examples of the presently disclosed technology are depicted.
[0020] Figure 6 depicts example operations for training a machine learning model to classify vehicle tampering events based on audio data, according to various examples of the presently disclosed technology.
[0021] Figure 7 depicts additional example operations for training a machine learning model to classify vehicle tampering events based on audio data, according to various examples of the presently disclosed technology.
[0022] Figure 8 depicts example acceleration-related data and audio data from an imitation / simulated catalytic converter theft event, according to various examples of the presently disclosed technology.
[0023] Figure 9 depicts a block diagram of an example computer system in which various examples described herein may be implemented.
[0024] The drawings are not exhaustive and do not limit the present disclosure to the precise forms disclosed. DETAILED DESCRIPTION
[0025] Examples of the presently disclosed technology are described in more detail in conjunction with the following figures.
[0026] Figure 1 illustrates the architecture of an example alarm system 100, according to various examples of the presently disclosed technology.
[0027] As depicted, alarm system 100 includes sensor 110, alarm system 120, and control unit 130. Sensor 110 and alarm system 120 may communicate with control unit 130 via communication circuitry 132 (described in more detail below). Although sensor 110 and alarm system 120 are depicted as communicating with control unit 130, they may also communicate with each other and with the external world via wireless or wired communication. Although depicted as a single control unit, control unit 130 may be implemented via multiple control units or as part of a control unit.
[0028] Sensor 110 may include various types of sensors for detecting objects in the environment of a vehicle (e.g., potential burglars / criminals, tools used by potential burglars / criminals, etc.). For example, sensor 110 may include an accelerometer 112 (in various examples, multiple accelerometers may be included in sensor 110), an audio sensor 114 (in various examples, multiple audio sensors may be included in sensor 110), and other sensors 116. Sensor 110 may include any combination of only data collection sensors and processing sensors, where the only data collection sensors provide only raw data to the control unit 130, and the processing sensors process the raw data (e.g., raw audio data, raw acceleration data, raw jerk data, etc.) and provide processed data (e.g., processed audio data, processed acceleration data, processed jerk data, etc.) to the control unit 130. Some of these sensors may provide a combination of raw data and processed data to the control unit 130.
[0029] As described above, accelerometer 112 may include means for measuring acceleration-related data of the body to which the accelerometer is mounted. The acceleration-related data may include the relative acceleration of the body (i.e., the relative rate of change of velocity) and / or relative jerk (i.e., the relative rate of change of acceleration). For example, if accelerometer 112 is mounted to a partition of a van, accelerometer 112 may measure the relative acceleration and / or relative jerk of the partition when the partition moves / is disturbed (such as during a vehicle tampering event). Correlatively, if accelerometer 112 is mounted to a door or wall of a vehicle, accelerometer 112 may measure the relative acceleration and / or relative jerk of the door / wall when the door / wall moves / is disturbed (such as during a vehicle tampering event). Since examples of the presently disclosed technology are designed with understanding, certain bodies of a vehicle may move / vibrate in a similar resonance as other bodies of the vehicle that are directly disturbed during a vehicle tampering event. For example, the floor or wall of a vehicle may move / vibrate in a similar resonance as an exhaust system mounted to the underside of the vehicle. Thus, even if accelerometer 112 is mounted to the floor or wall of a vehicle, it may effectively detect acceleration-related data (i.e., relative acceleration or relative jerk) associated with a burglar cutting / sawing a catalytic converter from the underside of the vehicle.
[0030] Audio sensor 114 may include various types of audio sensors, including microphones. In various examples, audio sensor 114 may be located / mounted within the interior space of a vehicle (e.g., the loading space of a van). In certain examples, audio sensor 114 may be strategically positioned / directed such that it can detect sound waves with reduced background noise.
[0031] As described above, examples of the presently disclosed techniques utilize accelerometers and audio sensors because they can detect vehicle tampering events without relying on a direct / clear line of sight to the vehicle tampering event. In contrast, other types of sensors utilized by many existing alarm systems rely on a clear line of sight to detect vehicle tampering events. Examples of these “line-of-sight-dependent” sensors include cameras (e.g., a dashboard camera of a vehicle), computer vision, proximity sensors (e.g., radar, lidar, and / or sonar sensors of a vehicle), passive infrared sensors, and the like.
[0032] Nevertheless, in addition to other types of sensors, other sensors 116 may also include the above-referenced “line-of-sight-dependent” sensors. Such sensors may still be useful for detecting certain vehicle tampering events where a clear / direct line of sight to the vehicle tampering event (or the consequences of the vehicle tampering event) is available. By way of example, other sensors 116 may include vibration sensors, motion sensors, temperature sensors, door lock sensors, tilt sensors, wireless signal detectors, GPS devices, cameras with hardware and software for performing face or object recognition, lidar, radar, ANPR, radio frequency detectors, Bluetooth detectors, gyroscopes, passive infrared (PIR) detectors, detectors that determine the identity of a mobile device (i.e., a mobile phone), or any combination thereof.
[0033] In some examples, sensors 110 may be encapsulated within a sensing unit of alarm system 100. Such a sensing unit may be mounted to or within a protected vehicle (e.g., within a cargo space or other interior space of the vehicle). However, in other examples, sensors 110 may be encapsulated / mounted independently of one another. In various examples, one or more of sensors 110 may be sensors of the protected vehicle / vehicle system. However, in other examples, sensors 110 may be implemented independently of the vehicle / vehicle system.
[0034] As described above, in some examples, accelerometer 112 and audio sensor 114 may operate continuously, or at least continuously when the vehicle is unattended. In contrast, some of the other sensors 116 (e.g., cameras, radar, lidar, etc.) may be activated only after a vehicle tampering event has been classified and verified. Activating these other sensors may help detect and / or identify the burglar / suspect. Relatedly, by activating these sensors only in response to classifying and verifying a vehicle tampering event, alarm system 100 may conserve power / reduce power consumption.
[0035] Now referring to control unit 130, control unit 130 can receive any combination of raw data and processed data from sensor 110. Based on the received data, control unit 130 can classify and verify vehicle tampering events and put alarm system 100 on high alert. For example, control unit 130 can: (1) provide audio data received from audio sensor 114 to a machine learning model trained with audio signatures of known tampering events; (2) in response to the machine learning model classifying a potential vehicle tampering event that generated the audio data as a vehicle tampering event based on the audio data, compare acceleration-related data received from accelerometer 112 during the potential vehicle tampering event with a threshold; and (3) in response to determining that the acceleration-related data exceeds the threshold, put alarm system 100 on high alert based on the vehicle tampering event classification. As described above, putting alarm system 100 on high alert can include any combination of the following: (a) sending instructions to activate one or more alarm systems of alarm system 120; and (b) sending instructions to activate one or more sensors in other sensors 116 (e.g., cameras). As described above, in various examples, control unit 130 can process data received from sensor 110 into a time format more suitable for use by a machine learning model. For example, control unit 130 can: (1) receive first audio data (e.g., raw or preprocessed audio data) from audio sensor 114; (2) encode the first audio data into a latent representation; and (3) divide the latent representation into time window frames and stack the time window frames to generate time audio data. However, in other examples, sensor 110 (e.g., audio sensor 114) can perform the above data encoding and latent representation time division / stacking.
[0036] As depicted, control unit 130 can include communication circuit 132, determination circuit 136, and power supply 139. Components within control unit 130 can communicate via a data bus and / or other suitable communication interfaces.
[0037] Communication circuit 132 can include at least one of wireless communication interface 133 (e.g., a transceiver with an antenna) and wired communication interface 134 (e.g., an I / O interface with an associated hardwired data port). Control unit 130 can utilize communication circuit 132 to communicate with sensor 110 and alarm system 120. Control unit 130 can also utilize communication circuit 132 to communicate with devices remote from alarm system 100 (e.g., vehicles, external alarm / surveillance systems, relevant authorities, etc.).
[0038] The wireless communication interface 133 may include a transceiver (i.e., a receiver and a transmitter) to allow wireless communication via various communication protocols (such as, WiFi, Zigbee, Bluetooth, Near Field Communication, etc.). As described above, the wireless communication interface 133 may include an antenna that is coupled to the transceiver to wirelessly transmit and receive radio signals. These radio signals may include information sent to and from the sensor 110 and the alarm system 120. These radio signals may also include radio signals sent to and from devices away from the alarm system 100 (e.g., vehicles, external alarm / surveillance systems, relevant authorities, etc.).
[0039] The wired communication interface 134 may include a receiver and a transmitter for hardwired communication with other components of the alarm system 100 (e.g., the sensor 110 and the alarm system 120). For example, the wired communication interface 134 may provide a hardwired interface to other components including the sensor 110 and the alarm system 120. The wired communication interface 134 may communicate with these components using Ethernet or any number of other wired communication protocols. In various examples, the wired communication interface 134 may also communicate with a device away from the alarm system 100 (e.g., a vehicle in which the alarm system 100 is located / installed).
[0040] As depicted, the determination circuit 136 includes a processor 137 and a memory 138. The processor 137 may include one or more processing resources, such as a GPU, a CPU, a microprocessor, etc.
[0041] The memory 138 may include one or more modules of various forms of memory / data storage devices (e.g., flash memory, RAM, etc.) for storing various data, parameters, and operation instructions utilized by the processor 137 and any other suitable information. For example, the memory 138 may store audio and / or time characteristics of known vehicle tampering events, acceleration-related thresholds or characteristics for verifying vehicle tampering event classification, etc.
[0042] Although specific examples of Figure 1 are shown using processor and memory circuitry, any form of circuitry (including, for example, hardware, software, or a combination thereof) may be utilized to implement the determination circuit 136. As another example, one or more processors, controllers, ASICs, PLAs, PALs, CPLDs, FPGAs, logic components, software routines, or other mechanisms may be used to implement the control unit 130.
[0043] Power supply 139 may include any suitable type of power supply. For example, power supply 139 may include one or more batteries (e.g., rechargeable or primary batteries including lithium-ion, lithium polymer, NiMH, NiCd, NiZn, NiH2, etc.), power connectors (e.g., for connection to vehicle-supplied power), and energy collectors (e.g., solar cell cores, piezoelectric systems, etc.).
[0044] Now referring to alarm system 120, as depicted, alarm system 120 may include various types of alarm systems, including visual alarm 122 (e.g., strobe light), audio alarm 124 (e.g., horn, siren, or other audible warning signal), external alarm notification system 126, and other alarms 128. As the name implies, external alarm notification system 126 may use wireless or wired communication to send alarm notifications to remote entities (e.g., relevant authorities, vehicles, etc.) in the same / similar manner as described for communication circuit 132 coupled to control unit 130. In some examples, one or more of the alarm systems 120 may be connected to or otherwise associated with a vehicle / vehicle system. For example, audio alarm 124 may be connected to a vehicle horn, or visual alarm may be connected to vehicle lights. However, in other embodiments, alarm system 120 may be implemented independently of the vehicle. Although not depicted, alarm system 120 may include its own processing resources and memory. The same may be true for sensor 110.
[0045] As described above, control unit 130 may place alarm system 100 on high alert by sending instructions to or otherwise activating one or more of the alarm systems 120.
[0046] Figure 2 An example method for vehicle tampering event classification and verification according to various examples of the presently disclosed technology is depicted. As depicted, the method may be performed by alarm system 200 of the presently disclosed technology. Alarm system 200 may be the same / similar to alarm system 100 described in connection with Figure 1 the description above.
[0047] In more detail, Figure 2Prior to reiterating, it is helpful to understand that examples of the presently disclosed technology are designed for intelligent insights. That is, certain vehicle tampering events have unique time signatures as well as unique audio signatures. For example, a handle pull tampering event (and the associated audio from the handle pull tampering event) typically has a duration of a few seconds or less. In contrast, a catalytic converter theft tampering event (and the associated audio from the catalytic converter tampering event) typically has a duration of 30 seconds to one minute. Based on this insight, examples of the presently disclosed technology can improve the accuracy of audio-based classification by leveraging a temporal convolutional network (TCN) model (e.g., Figure 2 the TCN audio model 206) that is particularly suitable for learning the temporal characteristics as well as the audio characteristics of vehicle tampering events. Correlatively, examples of the presently disclosed technology can convert audio data (e.g., raw or preprocessed audio data) into a temporal format that is more effectively / efficiently processed by the TCN model.
[0048] As depicted, the alarm system 200 can use an encoder 202 to encode audio data 201 (e.g., raw audio data or preprocessed audio data) into a latent representation 203. The alarm system 200 can then divide the latent representation 203 into time window frames and stack the time window frames to generate temporal audio data 205. The alarm system 200 can then provide the temporal audio data 205 to the TCN audio model 206 to classify a potential / suspected vehicle tampering event (or more precisely, the noise event from which the audio data 201 is derived) that produced the audio data 201. As described above, the TCN audio model 206 can be trained to classify potential / suspected vehicle tampering events using the audio / temporal audio characteristics of known vehicle tampering events.
[0049] As described above, during a potential / suspected vehicle tampering event, the audio data 201 can be acquired by an audio sensor of the alarm system 200. The audio data 201 can include raw audio data or preprocessed audio data (e.g., noise-removed purified audio data, labeled audio data, etc.). In many cases, the audio data 201 will capture features / information that are not detectable by the human ear from the audio signal.
[0050] The latent representation 203 can include a lower-dimensional representation of the audio data 201 that captures the key / important features of the audio data 201. As described above, the alarm system 200 can divide the latent representation 203 into time window frames and stack the time window frames to generate temporal audio data.
[0051] Examples of the latent representation 203 can include the root mean square, which for each time window frame includes the sum of the squared values of all samples in the time window frame divided by the number of samples in the corresponding time window frame. Another example of the latent representation 203 can include the amplitude envelope, which for each time window frame includes the maximum value in the time window frame. The latent representation 203 can also include the zero crossing rate, which represents the rate at which the audio signal (captured by the audio data 201) intersects the horizontal amplitude axis.
[0052] In additional examples, the latent representation 203 can include a spectrogram and / or one or more types of values / features derived from the spectrogram. As used herein, a spectrogram can refer to a visual representation of the spectrum of a signal as the signal varies over time. Thus, the spectrogram of the audio data 201 can include a visual representation of the spectrum of the audio signal captured within the audio data 201. By rapidly generating and analyzing the spectrogram (or other types of latent representations), the alert system 200 can classify vehicle tampering events based on audio signal information that is typically not detectable by the human ear. Additionally, using the TCN audio model 206, the alert system 200 can rapidly analyze the spectrogram (or other types of latent representations) to classify vehicle tampering events in real time (or near real time). Such rapid analysis can be crucial for detecting vehicle tampering events in sufficient time to take effective deterrent measures.
[0053] Examples of spectrogram values of the latent representation 203 can include: (a) spectral flux (e.g., the Euclidean distance of successive normalized spectra); (b) spectral centroid (e.g., where each frame of the amplitude spectrogram is normalized and treated as a distribution over frequency bins, from which the mean (centroid) is extracted for each frame); (c) spectral spread (e.g., the second central moment of the spectrum / the deviation of the spectrum from the spectral centroid); and (d) spectral roll-off (e.g., the frequency of each frame where a percentage of the total spectral energy is below a roll-off threshold). Another example can be spectral contrast, where each frame of the amplitude spectrogram is divided into subbands. For each subband, the energy contrast is estimated by comparing the average energy of the top quantile (peak energy) with the average energy of the bottom quantile (valley energy). High contrast values typically correspond to clear narrowband signals, while low contrast values correspond to broadband noise.
[0054] Referring again to Figure 2 , as depicted, the alert system 200 can provide the vehicle tampering event classification 209 from the TCN audio model 206 to the alert decision module 212. The output of the alert decision module 212 is based on the vehicle tampering event classification 209 and the threshold comparison 210 performed by the acceleration-related model 208.
[0055] As described above, the alarm system 200 can verify audio-based classification by analyzing acceleration-related data (e.g., relative acceleration data, relative jerk data, etc.) from the vehicle body to determine suspicious movement of the vehicle body during a potential / suspicious vehicle tampering event. Such acceleration-related data can be obtained by an accelerometer mounted to the vehicle body. This acceleration-related verification step can reduce the occurrence of audio-based false alarm classifications caused by other noise events near the vehicle that have audio characteristics similar to vehicle tampering events (e.g., drilling or other noises from a construction site, rain, etc.).
[0056] Accordingly, the alarm system 200 can provide acceleration-related data 207 to the acceleration-related model 208. As described above, the acceleration-related data 207 can include (raw or pre-processed) acceleration data and / or (raw or pre-processed) jerk data from the vehicle body during a potential / suspicious vehicle tampering event. As described above, the acceleration-related data 207 can be obtained by an accelerometer mounted to the vehicle body.
[0057] The acceleration-related model 208 can include a machine learning model trained to analyze acceleration-related data, or another process / algorithm / module capable of comparing the acceleration-related data with a corresponding threshold. Accordingly, the acceleration-related model 208 can compare the acceleration-related data 207 with the corresponding threshold and provide the threshold comparison 210 to the alarm decision module 212.
[0058] Accordingly, in response to the threshold comparison 210 indicating an exceedance of the threshold, the alarm decision module 212 can place the alarm system 200 on high alert based on the vehicle tampering event classification 206. In contrast, in response to the threshold comparison 210 indicating no exceedance of the threshold, the alarm decision module 212 can discard the vehicle tampering event classification 209 and maintain the default alert state of the alarm system 200.
[0059] To summarize and restate the above description in a slightly different way, encoding the audio data 201 into the latent representation 203 is the process of transforming the audio data 201 into a compact and informative form (i.e., the latent representation 203 and then the temporal audio data 205) that can be effectively used by a machine learning model such as the TCN audio model 206.
[0060] As described above, the first step in transforming the audio data 201 into the latent representation 203 is to extract features from the audio data 201 that capture relevant information for the identification task. The encoder 202 (or a separate feature extraction module or a separate convolutional layer dedicated to feature extraction) can perform such feature extraction. Examples of the extracted features can include spectrograms, Mel-frequency cepstral coefficients (MFCCs), and chroma features.
[0061] As described above, a spectrogram visually represents how the spectral density of a signal changes over time, thus effectively capturing the time-varying frequency content.
[0062] MFCCs are coefficients that together form the mel-scale cepstral representation of an audio segment. They are derived from the Fourier transform of a signal, but are represented on the mel scale, which approximates the human ear's response to different frequencies.
[0063] The extracted features cited above may still be high-dimensional, which can be challenging for the effective processing of machine learning models (e.g., the TCN audio model 206). Thus, an example can utilize an autoencoder (e.g., the encoder 202) to reduce the dimensionality. The goal can be to retain the most informative aspects of the audio data 201 while reducing the overall size of the data that will be provided to the TCN audio model 206.
[0064] As described above, the encoder 202 can be used to map a high-dimensional feature space to a lower-dimensional latent space 203. The encoder 202 can be trained to produce a latent representation that retains as much relevant information as possible from the original audio data 201. This process can involve minimizing a loss function that measures the difference between the original audio data 201 and the reconstructed data (i.e., the latent representation 203), thus ensuring that the latent representation 203 is meaningful.
[0065] As described above, then, the alert system 200 can divide the latent representation 203 into time window frames and stack the time window frames to generate time audio data 205. Then, the alert system 200 can provide the time audio data 205 to the TCN audio model 206 to classify a potential / suspicious vehicle tampering event (or more precisely, the noise event from which the audio data 201 is derived) that produced the audio data 201.
[0066] In some embodiments, each convolutional layer of the TCN audio model 206 can be considered a feature transformer, thus further refining the time audio data 205 at each layer. The TCN audio model 206 can learn to identify patterns across both features and time series.
[0067] In some embodiments, the entire process depicted in Figure 2 can be learned end-to-end. This means that the above-described feature extraction, latent representation learning, and classification are all parts of a single model that is trained together. In some examples, a single model can be designed to include an encoder (e.g., the encoder 202) as its initial layer, which automatically learns to convert the audio data 201 into a latent representation 203 as part of its training process.
[0068] Figure 3Depicts example operations for performing vehicle tampering event classification and verification according to various examples of the presently disclosed technology. As depicted, Figure 3 the operations may be performed by the alarm system 300 of the presently disclosed technology. The alarm system 300 may be the same / similar to the alarm system 100 described in conjunction with Figure 1 Figure 1.
[0069] As depicted, the alarm system 300 may perform operation 302 to provide audio data from a potential vehicle tampering event to a machine learning model trained with audio features of known vehicle tampering events. The machine learning model may then classify the potential vehicle tampering event as a vehicle tampering event based on the audio data. The vehicle tampering event classification may include any number of vehicle tampering event classifications, including: (a) handle pull tampering event classification; (b) drilling tampering event classification; (c) key-lock tampering event classification; (d) metal stripping related tampering event classification (i.e., vehicle tampering event classification related to the tearing or stripping of the metal body of the vehicle); and (e) catalytic converter theft tampering event classification.
[0070] In some embodiments, the audio data may include temporal audio data (e.g., audio data including stacked temporal window frames). Correlatively, the machine learning model may include a temporal convolutional network (TCN) model. In these examples, the alarm system 300 may perform further operations to: (a) receive first audio data from a potential vehicle tampering event; (b) encode the first audio data into a latent representation (i.e., a lower-dimensional representation of the first audio data that captures the important / critical features of the first audio data); (c) divide the latent representation into temporal window frames and stack the temporal window frames to generate temporal audio data; and (d) provide the temporal audio data to the TCN model. The first audio data may include raw audio data or preprocessed audio data and may be received from an audio sensor of the alarm system 300. The audio sensor may be located within the interior space of a vehicle protected by the alarm system 300.
[0071] In response to the machine learning model classifying a potential vehicle tampering event as a vehicle tampering event (i.e., based on audio data), the alarm system 300 may perform operation 304 to compare acceleration-related data from the vehicle's body (e.g., the body of the vehicle protected by the alarm system 300) during the potential vehicle tampering event with a threshold. As described above, the acceleration-related data may include at least one of relative acceleration data from the vehicle's body and relative jerk data from the vehicle's body. Correspondingly, the threshold may include at least one of an acceleration threshold and a jerk threshold. The acceleration-related data may include raw or preprocessed data and may be received from an accelerometer of the alarm system 300 installed on the surface of the vehicle's body.
[0072] In response to determining that the acceleration-related data exceeds the threshold, the alarm system 300 may perform operation 306(a) to place the alarm system 300 on high alert based on the vehicle tampering event classification. Placing the alarm system 300 on high alert may include at least one of the following: (a) activating additional sensors of the alarm system 300 (e.g., a camera of the alarm system 300); (b) activating an audio alarm; (c) activating a visual alarm; and (d) sending an alarm notification to a location remote from the alarm system 300 (e.g., sending an alarm to a separate monitoring system, sending an alarm to relevant authorities, etc.). As described above, based on the vehicle tampering event classification, what constitutes placing the alarm system 300 may be different. For example, the alarm system 300 may be placed in a first high alert state in response to a handle pull tampering event classification and in a second high alert state in response to a catalytic converter theft tampering event classification.
[0073] In contrast to the above paragraphs, in response to determining that the acceleration-related data does not exceed the threshold, the alarm system 300 may perform operation 306(b) to discard the vehicle tampering event classification and maintain the default alert state of the alarm system 300.
[0074] As described above, the acceleration-related verification steps of operations 304 and 306(a) / (b) may reduce the occurrence of audio-based false alarm classifications caused by other noise events near the vehicle that have audio characteristics similar to vehicle tampering events (e.g., drilling or other noises from a construction site, rain, etc.). Reducing the occurrence of false alarm classifications has many advantages, including: (a) increasing consumer trust in the alarm system 300; (b) reducing the annoyance of false alarms; (c) saving power in embodiments where additional parts of the alarm system 300 are awakened / activated in response to a verified detection / classification of a vehicle tampering event; and so on.
[0075] Figure 4Depicts example operations for classifying vehicle tampering events according to various examples of the presently disclosed technology. As depicted, Figure 4 the operations may be performed by the alarm system 400 of the presently disclosed technology. The alarm system 400 may be the same / similar to the alarm system 100 described in conjunction with Figure 1 description.
[0076] As depicted, the alarm system 400 may perform operation 402 to receive audio data from a potential vehicle tampering event. The audio data may be raw or pre - processed data and may be received from an audio sensor of the alarm system 400.
[0077] The alarm system 400 may perform operation 404 to encode the audio data into a latent representation. The latent representation may include a lower - dimensional representation of the audio data that captures key / important features of the audio data. Examples of the latent representation of audio data are described in more detail in conjunction with Figure 2 description.
[0078] The alarm system 400 may perform operation 406 to divide the latent representation into time - window frames and stack the time - window frames to generate time - audio data.
[0079] Based on the time - audio data, the alarm system 400 may perform operation 408 to classify the potential vehicle tampering event as a catalytic converter theft tampering event using a trained temporal convolutional network (TCN) model.
[0080] Then, the alarm system 400 may perform operation 410 to place the alarm system 400 on high alert based on the catalytic converter theft tampering event classification. As described above, placing the alarm system 400 on high alert may include at least one of the following: (a) activating additional sensors of the alarm system 400 (e.g., a camera of the alarm system 400); (b) activating an audio alarm; (c) activating a visual alarm; and (d) sending an alarm notification to a location remote from the alarm system 400 (e.g., sending an alarm to a separate monitoring system, sending an alarm to relevant authorities, etc.).
[0081] In certain embodiments, the alarm system 400 may perform further operations to compare acceleration - related data of the movement of a body of the vehicle (e.g., the vehicle protected by the alarm system 400) during a potential vehicle tampering event with a threshold. In these embodiments, placing the alarm system 400 on high alert based on the catalytic converter theft tampering event classification may include placing the alarm system 400 on high alert based on the catalytic converter theft tampering event classification in response to determining that the acceleration - related data exceeds the threshold.
[0082] As described above, by using temporal audio data and a TCN model, the alarm system 400 can improve classification accuracy relative to alternative methods lacking such a temporal approach. This improvement in classification accuracy has several advantages, including: (a) increasing consumer trust; (b) reducing the annoyance of false alarms; (c) saving power in implementations where additional portions of the alarm system 400 are awakened / activated in response to a verified detection / classification of a vehicle tampering event; (d) allowing for customized responses to specific types of vehicle tampering events; and so on.
[0083] Figure 5 Depicts example operations for classifying tampering events involving personal property based on audio data, according to various examples of the presently disclosed technology. As depicted, Figure 5 the operations may be performed by an alarm system 500 of the presently disclosed technology. The alarm system 500 may be the same / similar to the alarm system 100 described in connection with Figure 1 the description.
[0084] As described above, it should be understood that the principles disclosed herein may be applied to property types other than vehicles. For example, certain types of personal property may be stored in an enclosed space where the line of sight to the personal property and / or potential tampering events will be blocked or obstructed. As depicted, the alarm system 500 may perform operations to protect such personal property (which may or may not be a vehicle).
[0085] As depicted, the alarm system 500 may perform operation 502 to provide audio data from a potential tampering event to a machine learning model trained with audio characteristics of known tampering events. The machine learning model may then classify the potential tampering event as a tampering event based on the audio data. The tampering event classification may include any number of tampering event classifications, including: (a) handle pull or door pull tampering event classification; (b) drilling tampering event classification; (c) key-lock tampering event classification; (d) metal stripping-related tampering event classification (i.e., a tampering event classification related to the tearing or stripping of a metal body); and so on.
[0086] In some implementations, the machine learning model may classify the potential tampering event as a (non-tampering) event. These (non-tampering) event classifications may then be used to inform later classifications. For example, the machine learning model may classify the potential tampering event as a lock / unlock event. Here, the classification of a lock / unlock event may indicate to the alarm system 500 that the immediately following event is less likely to be related to a tampering event (since it is likely that an authorized individual is more likely to be involved in a lock / unlock event).
[0087] In some embodiments, the audio data may include temporal audio data (e.g., audio data including stacked temporal window frames). Correspondingly, the machine learning model may include a temporal convolutional network (TCN) model. In these examples, the alarm system 500 may perform further operations to: (a) receive first audio data from a potential tampering event; (b) encode the first audio data into a latent representation (i.e., a lower-dimensional representation of the first audio data that captures the important / critical features of the first audio data); (c) divide the latent representation into temporal window frames and stack the temporal window frames to generate temporal audio data; and (d) provide the temporal audio data to the TCN model. The first audio data may include raw audio data or preprocessed audio data and may be received from an audio sensor of the alarm system 500. The audio sensor may be located within the interior space of the property protected by the alarm system 500.
[0088] In response to the machine learning model classifying the potential tampering event as a tampering event (i.e., based on the audio data), the alarm system 500 may perform operation 504 to compare acceleration-related data of a subject of the property (e.g., the property protected by the alarm system 500) during the potential tampering event with a threshold. As described above, the acceleration-related data may include at least one of relative acceleration data of the subject of the property and relative jerk data of the subject of the property. Correspondingly, the threshold may include at least one of an acceleration threshold and a jerk threshold. The acceleration-related data may include raw or preprocessed data and may be received from an accelerometer of the alarm system 500 installed on the surface of the subject of the property.
[0089] In response to determining that the acceleration-related data exceeds the threshold, the alarm system 500 may perform operation 506(a) to place the alarm system 500 in a heightened state of alert based on the tampering event classification. Placing the alarm system 500 in a heightened state of alert may include at least one of the following: (a) activating additional sensors of the alarm system 500 (e.g., a camera of the alarm system 500); (b) activating an audio alarm; (c) activating a visual alarm; and (d) sending an alarm notification to a location remote from the alarm system 500 (e.g., sending an alarm to a separate monitoring system, sending an alarm to relevant authorities, etc.). As described above, based on the vehicle tampering event classification, what constitutes placing the alarm system 500 may be different. For example, the alarm system 500 may be placed in a first heightened state of alert in response to a handle pull tampering event classification and in a second heightened state of alert in response to a metal stripping-related tampering event classification.
[0090] Conversely, in response to determining that the acceleration-related data does not exceed the threshold, the alarm system 500 may perform operation 506(b) to discard the tampering event classification and maintain the default armed state of the alarm system 500.
[0091] As described above, the acceleration-related verification steps of operations 504 and 506(a) / (b) can reduce the occurrence of audio-based false alarm classifications caused by other noise events near the vehicle that have audio characteristics similar to vehicle tampering events (e.g., drilling or other noises from a construction site, rain, etc.). Reducing the occurrence of false alarm classifications has many advantages, including: (a) increasing consumer trust in the alarm system 500; (b) reducing the annoyance of false alarms; (c) saving power in implementations where additional parts of the alarm system 500 are awakened / activated in response to a verified detection / classification of a vehicle tampering event; and so on.
[0092] Figure 6 Depicted are example operations for training a machine learning model to classify vehicle tampering events based on audio data according to various examples of the presently disclosed technology. As depicted, Figure 6 the operations may be performed by a training system 600 of the presently disclosed technology.
[0093] As depicted, the training system 600 may perform operation 602 to receive audio data from known catalytic converter theft tampering events. The received audio data may include raw audio data and / or preprocessed audio data.
[0094] The known catalytic converter theft tampering events may include mimicking / simulating catalytic converter theft performed on a vehicle in a controlled scenario. For example, during the mimicking / simulating catalytic converter theft, audio may be recorded while the catalytic converter is cut off from the vehicle's exhaust system. The known catalytic converter theft tampering events may also include actual catalytic converter thefts implemented "in the wild" (i.e., implemented catalytic converter theft tampering events) detected / classified by alarm systems deployed by the presently disclosed technology.
[0095] To improve machine learning model training, the training system 600 may perform operation 604 to process the audio data. Processing the audio data may include various types of processing, including any combination of the following: (a) purifying the audio data to remove foreign noise or otherwise discard noise data, (b) labeling the audio data as related to a catalytic converter theft tampering event, (c) converting the audio data to a latent representation, and (d) preparing temporal audio data by dividing the latent representation of the audio data into time window frames and stacking the time window frames to generate the temporal audio data.
[0096] Accordingly, the training system 600 can perform operation 606 to use the processed audio data to train a machine learning model (e.g., a TCN model) to classify potential vehicle tampering events as catalytic converter theft tampering events.
[0097] The training system 600 can also utilize audio data from known non-catalytic converter theft tampering events (e.g., handle pull tampering events, drilling tampering events, key-lock tampering events, metal stripping related tampering events, etc.) and known non-tampering events (e.g., construction site noise, rain, etc.) to train a machine learning model to classify different types of vehicle tampering events and distinguish them from non-tampering events. The audio data from known non-catalytic converter theft tampering events and known non-tampering events can be collected during mimic / simulate events in a laboratory scenario (e.g., mimic / simulate handle pull tampering events, mimic / simulate metal stripping related tampering events, mimic / simulate non-tampering events (such as operating a jackhammer on a section of road near a vehicle, etc.)). Such audio data can be processed in the same / similar manner as described in connection with operation 604 before being used as training data.
[0098] Figure 7 Additional example operations for training a machine learning model to classify vehicle tampering events based on audio data according to various examples of the presently disclosed technology are depicted. As depicted, Figure 7 the operations of can be performed by the training system 700 of the presently disclosed technology.
[0099] As depicted, the training system 700 can perform operation 702 to receive audio data from known catalytic converter theft tampering events. The received audio data can include raw audio data and / or preprocessed audio data.
[0100] Known catalytic converter theft tampering events can include mimic / simulate catalytic converter thefts performed on vehicles in a controlled scenario. For example, during a mimic / simulate catalytic converter theft, audio can be recorded while the catalytic converter is cut off from the vehicle's exhaust system. Known catalytic converter theft tampering events can also include actual catalytic converter thefts "in the wild" detected / classified by an alarm system deployed by the presently disclosed technology.
[0101] The training system 700 can perform operation 704 to encode the audio data into a latent representation. As described above, the latent representation of the respective audio data from respective known catalytic converter theft tampering events can include a lower-dimensional representation of the respective audio data that captures important / critical features of the respective audio data. Examples of latent representations are described in more detail in connection with Figure 2 More detailed examples of latent representations are described.
[0102] The training system 700 may perform operation 706 to divide the latent representation into time window frames and stack the time window frames (for the corresponding latent representation) to generate time audio data (for the corresponding known catalytic converter theft tampering event).
[0103] Then, the training system 700 may perform operation 708 to use the time audio data to train a temporal convolutional network (TCN) model to classify latent vehicle tampering events as catalytic converter theft tampering events.
[0104] As described above, the training system 700 may also utilize audio data from known non-catalytic converter theft tampering events (e.g., handle pull tampering events, drilling tampering events, key-lock tampering events, metal stripping related tampering events, etc.) and known non-tampering events (e.g., construction site noise, rain, etc.) to train the TCN model to classify different types of vehicle tampering events and distinguish them from non-tampering events. The audio data from known non-catalytic converter theft tampering events and known non-tampering events may be collected during mimic / simulate events in a laboratory scenario (e.g., mimic / simulate handle pull tampering events, mimic / simulate metal stripping related tampering events, mimic / simulate non-tampering events (such as operating a jackhammer near a vehicle, etc.)). Such audio data may be processed in the same / similar manner as described in connection with operations 704 to 706 before being used as training data.
[0105] Figure 8 Example acceleration-related data and audio data from mimic / simulate catalytic converter theft events according to various examples of the presently disclosed technology are depicted.
[0106] That is, graphs 802 to 806 show acceleration-related data from mimic / simulate catalytic converter theft events. Graphs 822 to 826 show audio data from mimic / simulate catalytic converter theft events. As depicted, all six graphs use a common time scale along their respective "x" axes.
[0107] As described above, the mimicking / emulating catalytic converter theft event includes an individual cutting off the catalytic converter from the underside of the vehicle. Three accelerometers each mounted to the body of the vehicle record acceleration-related data from the mimicking / emulating catalytic converter theft event. That is, a first accelerometer mounted to the body within the front passenger compartment space of the vehicle records the acceleration-related data depicted with a first shading in graphs 802 to 806. A second accelerometer mounted to the body within the middle passenger compartment space of the vehicle records the acceleration-related data depicted with a second shading in graphs 802 to 806. A third accelerometer mounted to the body within the rear passenger compartment space of the vehicle records the acceleration-related data depicted with a third shading in graphs 802 to 806. Similarly, a first audio sensor (e.g., a first microphone) mounted within the front passenger compartment space of the vehicle records the audio data depicted in graph 822. A second audio sensor (e.g., a second microphone) mounted within the middle passenger compartment space of the vehicle records the audio data depicted in graph 824. A third audio sensor (e.g., a third microphone) mounted within the rear passenger compartment space of the vehicle records the audio data depicted in graph 826.
[0108] As depicted, in graphs 802 to 826, the mimicking / emulating catalytic converter theft event begins at approximately time t1 and ends at approximately time t2. Before t1 and after t2, the accelerometers and audio sensors record ambient vibrations / movements and ambient noise.
[0109] Graph 802 depicts a histogram of the band-pass filtered jerk (i.e., a canonical example of acceleration-related data) over time during the mimicking / emulating catalytic converter theft event. Here, the band-pass filtered jerk can be associated with all three x-, y-, and z-axes in space. As described above, three different shading lines (better seen in graph 806) represent the band-pass filtered jerk data obtained from three differently positioned accelerometers. In some embodiments, the band-pass filtered jerk data can be obtained directly from the accelerometer, while in other embodiments, the band-pass filtered jerk data can be a post-processed version of the acceleration-related data obtained from the accelerometer.
[0110] Graph 804 depicts a histogram of jerk (i.e., a canonical example of acceleration-related data) over time during an emulated / simulated catalytic converter theft event. As described above, the jerk data can be associated with all three x-, y-, and z-axes in space. Similarly, three different shaded lines (seen better in Graph 806) represent jerk data obtained from three differently positioned accelerometers. In some embodiments, the jerk data can be obtained directly from the accelerometers, while in other embodiments, the jerk data can be a post-processed version of the acceleration-related data obtained from the accelerometers.
[0111] Graph 806 depicts a histogram of the root mean square (RMS) of jerk (i.e., a canonical example of acceleration-related data) over time during an emulated / simulated catalytic converter theft event. As described above, the RMS jerk data can be associated with all three x-, y-, and z-axes in space. Similarly, three different shaded lines represent RMS jerk data obtained from three differently positioned accelerometers. In some embodiments, the RMS jerk data can be obtained directly from the accelerometers, while in other embodiments, the RMS jerk data can be a post-processed version of the acceleration-related data obtained from the accelerometers.
[0112] Graphs 822 through 826 each depict a histogram of audio data from an emulated / simulated catalytic converter theft event. That is, Graph 822 depicts a spectrogram of audio data obtained from a first audio sensor, Graph 824 depicts a spectrogram of audio data obtained from a second audio sensor, and Graph 826 depicts a spectrogram of audio data obtained from a third audio sensor.
[0113] As depicted, the emulated / simulated catalytic converter theft event produces a unique temporal audio signature between times t1 and t2, which is visualized in the spectrograms of Graphs 822 through 826. Correlatively, the emulated / simulated catalytic converter theft event produces distinct acceleration-related signatures during this time period. These signatures can be contrasted with the times shown before t1 and after t2.
[0114] As described throughout this application, examples can detect / classify catalytic converter tampering events based on the associated unique audio signatures of catalytic converter tampering events (and other vehicle tampering events). Additionally, examples can validate these audio-based classifications by analyzing acceleration-related data (see, e.g., Graphs 802 through 806) to determine suspicious movement of the body of the vehicle during a potential / suspected vehicle tampering event.
[0115] From Figure 8It can also be seen that the curves 822 to 826 are substantially the same. This can indicate that in some embodiments, the audio sensor placement can have a reduced impact on the acquired / analyzed audio data. In a specific example related to Figure 8 simulated / catalytic converter theft events associated with, the vehicle is a van, and three audio sensors are installed inside the van's compartment. Here, the similarity between the curves 822 to 826 can indicate that the body of the van effectively acts as an amplifier that improves the audio sensing quality / consistency. Such insights can be utilized when classifying vehicle tampering events based on audio data.
[0116] Figure 9 A block diagram of an example computer system 900 in which various examples described herein can be implemented is depicted. For example, the computer system 900 can be used to implement Figure 1 the alarm system 100 of Figure 2 the alarm system 200 of Figure 3 the alarm system 300 of Figure 4 the alarm system 400 of Figure 5 the alarm system 500 of Figure 6 the training system 600 of Figure 7 and the training system 700 of. The computer system 900 includes a bus 902 or other communication mechanism for conveying information, and one or more hardware processors 904 coupled to the bus 902 for processing information. The hardware processor 904 can be, for example, one or more general-purpose microprocessors.
[0117] The computer system 900 also includes a main memory 906 coupled to the bus 902, such as a random access memory (RAM), cache, and / or other dynamic storage devices, for storing information and instructions to be executed by the processor 904. The main memory 906 can also be used to store temporary variables or other intermediate information during the execution of instructions by the processor 904. Such instructions, when stored in a storage medium accessible by the processor 904, cause the computer system 900 to become a special-purpose machine customized to perform the operations specified in the instructions.
[0118] The computer system 900 also includes a read-only memory (ROM) 908 or other static storage device coupled to the bus 902 for storing static information and instructions for the processor 904. A storage device 910 (such as a magnetic disk, optical disk, or USB thumb drive (flash drive), etc.) is provided and coupled to the bus 902 for storing information and instructions.
[0119] The computer system 900 can be coupled via a bus 902 to a display 912, such as a liquid crystal display (LCD) (or touch screen), for displaying information to a computer user. An input device 914, including alphanumeric keys and other keys, is coupled to the bus 902 for transmitting information and command selections to the processor 904. Another type of user input device is a cursor control 916, such as a mouse, trackball, or cursor direction keys, for transmitting direction information and command selections to the processor 904 and for controlling cursor movement on the display 912. In some examples, the same direction information and command selections as cursor control can be implemented by receiving touches on the touch screen without a cursor.
[0120] The computing system 900 can include a user interface module for implementing a GUI, which can be stored as executable software code executed by the computing device in a mass storage device. For example, this module and other modules can include components, such as software components, object-oriented software components, class components, and task components, procedures, functions, attributes, programs, subroutines, program code segments, drivers, firmware, microcode, circuits, data, databases, data structures, tables, arrays, and variables.
[0121] In general, as used herein, words such as "component", "engine", "system", "database", "data storage device", etc. can refer to logic embodied in hardware or firmware, or to a collection of software instructions, possibly having entry points and exit points written in a programming language such as Java, C, or C++. Software components can be compiled and linked into an executable program, installed in a dynamic link library, or can be written in an interpreted programming language such as BASIC, Perl, or Python. It should be understood that software components can call from other components or themselves, and / or can be called in response to detected events or interrupts. Software components configured to execute on a computing device can be provided on a computer-readable medium, such as a compact disc, digital video disc, flash drive, magnetic disk, or any other tangible medium, or as a digital download (and can initially be stored in a compressed or installable format that requires installation, decompression, or decryption before execution). Such software code can be stored, in part or in whole, on the memory device of the executing computing device for execution by the computing device. Software instructions can be embedded in firmware, such as an EPROM. It should also be understood that hardware components can be composed of connected logic units, such as gates and flip-flops, and / or can be composed of programmable units, such as programmable gate arrays or processors.
[0122] The computer system 900 may implement the techniques described herein using custom hardwired logic, one or more ASICs or FPGAs, firmware, and / or program logic that, in combination with the computer system, cause the computer system 900 to be or program the computer system to be a special-purpose machine. According to one example, the techniques herein are performed by the computer system 900 in response to one or more sequences of one or more instructions contained in the main memory 906 being executed by the processor 904. Such instructions may be read into the main memory 906 from another storage medium, such as the storage device 910. Execution of the instruction sequence contained in the main memory 906 causes the processor 904 to perform the process steps described herein. In an alternative example, hardwired circuitry may be used in place of or in combination with software instructions.
[0123] As used herein, the term "non-transitory medium" and like terms refer to any medium that stores data and / or instructions that cause a machine to operate in a particular manner. Such non-transitory media may include non-volatile media and / or volatile media. Non-volatile media includes, for example, optical or magnetic disks, such as the storage device 910. Volatile media includes dynamic memory, such as the main memory 906. Common forms of non-transitory media include, for example, floppy disks, flexible disks, hard disks, solid state drives, magnetic tape, or any other magnetic data storage medium, CD-ROM, any other optical data storage medium, any physical medium with patterns of holes, RAM, PROM, and EPROM, flash EPROM, NVRAM, any other memory chip or cartridge, and network versions thereof.
[0124] A non-transitory medium is different from a transmission medium but may be used in combination with it. A transmission medium participates in transferring information between non-transitory media. For example, a transmission medium includes coaxial cables, copper wire, and fiber optics, including the wires that comprise the bus 902. Transmission media can also take the form of acoustic or light waves, such as those generated during radio wave and infrared data communications.
[0125] The computer system 900 also includes a communication interface 918 coupled to the bus 902. The network interface 918 provides a two-way data communication link with one or more network links connected to one or more local networks. For example, the communication interface 918 can be an Integrated Services Digital Network (ISDN) card, a cable modem, a satellite modem, or a modem for providing a data communication connection with a corresponding type of telephone line. As another example, the network interface 918 can be a Local Area Network (LAN) card to provide a data communication connection with a compatible LAN (or a WAN component communicating with a WAN). A wireless link can also be implemented. In any such implementation, the network interface 918 transmits and receives electrical, electromagnetic, or optical indicators carrying digital data streams representing various types of information.
[0126] Network links typically provide data communication to other data devices through one or more networks. For example, a network link can provide a connection to a host computer or data equipment operated by an Internet Service Provider (ISP) through a local network. The ISP in turn provides data communication services through the global packet data communication network now commonly referred to as the "Internet". Both the local network and the Internet use electrical, electromagnetic, or optical indicators carrying digital data streams. The indicators through the various networks and the indicators on the network link and through the communication interface 918 (which carry digital data to and from the computer system 900) are example forms of transmission media.
[0127] The computer system 900 can send messages and receive data, including program code, through the network, network link, and communication interface 918. In the Internet example, a server can transmit the requested code of an application through the Internet, the ISP, the local network, and the communication interface 918.
[0128] The received code can be executed by the processor 904 when it is received, and / or stored in the storage device 910 or other non-volatile storage devices for later execution.
[0129] Each of the processes, methods, and algorithms described in the foregoing sections can be embodied in code components executed by one or more computer systems or computer processors including computer hardware, and be fully or partially automated by the code components. One or more computer systems or computer processors can also operate to support the execution of related operations in a "cloud computing" environment or as "software as a service" (SaaS). The processes and algorithms can be implemented partially or fully in dedicated circuitry. The various features and processes described above can be used independently of each other or can be combined in various ways. Different combinations and sub - combinations are intended to fall within the scope of the present disclosure, and in some embodiments, certain method or process blocks can be omitted. The methods and processes described herein are also not limited to any particular sequence, and the blocks or states associated therewith can be executed in a suitable other sequence, or can be executed in parallel or in some other manner. Blocks or states can be added to or removed from the disclosed exemplary examples. The execution of certain operations or processes can be distributed among computer systems or computer processors that not only reside within a single machine but are also deployed across several machines.
[0130] As used herein, a circuit can be implemented using any form of hardware, software, or a combination thereof. For example, one or more processors, controllers, ASICs, PLAs, PALs, CPLDs, FPGAs, logic components, software routines, or other mechanisms can be implemented to constitute a circuit. In an embodiment, the various circuits described herein can be implemented as discrete circuits, or the described functions and features can be partially or fully shared among one or more circuits. Even though various features or functional elements can be described or claimed separately as separate circuits, these features and functionality can be shared among one or more common circuits, and such a description should not require or imply the need for separate circuits to implement such features or functionality. Where the circuit is implemented fully or partially using software, such software can be implemented to operate with a computing or processing system (such as computer system 900) capable of implementing the functionality described thereof.
[0131] As used herein, the term "or" can be interpreted in an inclusive or exclusive sense. Additionally, a description of a resource, operation, or structure in the singular form should not be construed as excluding the plural form. Unless otherwise expressly specified or otherwise understood within the context in use, conditional language such as, without limitation, "can," "could," "might," or "may" generally intends to convey that certain examples include certain features, elements, and / or steps, while other examples do not include certain features, elements, and / or steps.
[0132] Unless otherwise expressly stated, the terms and phrases used in this document and their variants shall be construed as open-ended and not limiting. Adjectives such as "conventional", "traditional", "normal", "standard", "known" and terms with similar meanings shall not be construed as restricting the items described to a given time period or to items available at a given time, but shall be construed to cover conventional, traditional, normal or standard techniques available or known at any time, present or future. In some instances, the presence of broad words and phrases such as "one or more", "at least", "but not limited to" or other similar phrases shall not be construed to mean that a narrower situation is contemplated or required in instances where such broad phrases may not be present.
[0133] According to the present invention, a method for protecting property includes: providing audio data from a potential tampering event to a machine learning model trained with audio characteristics of known tampering events; in response to the machine learning model classifying the potential tampering event as a tampering event, comparing acceleration-related data of a subject of the property during the potential tampering event with a threshold; and in response to determining that the acceleration-related data exceeds the threshold, placing an alarm system on high alert based on the tampering event classification.
[0134] In one aspect of the present invention, the method includes: in response to determining that the acceleration-related data does not exceed the threshold, discarding the tampering event classification and maintaining the default alert state of the alarm system.
[0135] In one aspect of the present invention, the audio data includes temporal audio data; and the machine learning model includes a temporal convolutional network (TCN) model.
[0136] In one aspect of the present invention, the method includes: receiving first audio data from the potential tampering event; encoding the first audio data into a latent representation; dividing the latent representation into time window frames and stacking the time window frames to generate the temporal audio data; and providing the temporal audio data to the TCN model.
[0137] In one aspect of the present invention, the method includes: receiving first audio data from the potential tampering event from an audio sensor of the alarm system; processing the first audio data to generate the audio data; and during the tampering event, receiving the acceleration-related data of the subject of the property from an accelerometer of the alarm system.
[0138] In one aspect of the present invention, the property includes a vehicle, and the body of the property includes the body of the vehicle; the audio sensor is located within the interior space of the vehicle; the accelerometer is mounted to the surface of the body of the vehicle; and the surface of the body of the vehicle interfaces with the interior space of the vehicle.
[0139] In one aspect of the present invention, the tampering event classification includes at least one of the following: handle pull tampering event classification; drilling tampering event classification; key-lock tampering event classification; metal stripping related tampering event classification; and catalytic converter theft tampering event classification.
[0140] In one aspect of the present invention, placing the alarm system in the heightened state of alert includes at least one of the following: activating an additional sensor of the alarm system; activating an audio alarm; activating a visual alarm; and sending an alarm notification to a location remote from the alarm system.
[0141] According to the present invention, there is provided an alarm system having: one or more processing resources; and a memory coupled to the one or more processing resources, the memory storing instructions executable by the one or more processing resources to: provide audio data from a potential tampering event involving a vehicle to a machine learning model trained with audio characteristics of known vehicle tampering events; in response to the machine learning model classifying the potential tampering event as a vehicle tampering event, compare acceleration-related data from the body of the vehicle during the potential tampering event with a threshold; in response to determining that the acceleration-related data exceeds the threshold, place the alarm system in a heightened state of alert based on the vehicle tampering event classification; and in response to determining that the acceleration-related data does not exceed the threshold, discard the vehicle tampering event classification and maintain the default state of alert of the alarm system.
[0142] According to an embodiment, the present invention is further characterized by an audio sensor communicating with the one or more processing resources, wherein: the audio sensor collects first audio data from the potential tampering event; and the audio data is derived from the first audio data.
[0143] According to an embodiment, the audio data includes temporal audio data; and the memory stores further instructions executable by the one or more processing resources to: receive the first audio data from the potential vehicle tampering event from the audio sensor; encode the first audio data into a latent representation; and divide the latent representation into time window frames and stack the time window frames to generate the temporal audio data.
[0144] According to an embodiment, the machine learning model includes a Temporal Convolutional Network (TCN) model.
[0145] According to an embodiment, the present invention is further characterized by an accelerometer in communication with the one or more processing resources, wherein: the accelerometer collects first acceleration-related data from the body of the vehicle during the potential tampering event; and the acceleration-related data is derived from the first acceleration-related data.
[0146] According to an embodiment, the audio sensor is located within the interior space of the vehicle; the accelerometer is mounted to the surface of the body of the vehicle; and the surface of the body of the vehicle interfaces with the interior space of the vehicle.
[0147] According to an embodiment, the vehicle tampering event classification includes at least one of the following: handle pull tampering event classification; drilling tampering event classification; key-lock tampering event classification; metal stripping-related tampering event classification; and catalytic converter theft tampering event classification.
[0148] According to the present invention, a non-transitory computer-readable medium storing instructions that, when executed by one or more processors, cause the one or more processors to: provide time audio data from a potential tampering event involving a vehicle to a Temporal Convolutional Network (TCN) model trained with time audio features of known vehicle tampering events; in response to the TCN model classifying the potential tampering event as a vehicle tampering event, compare acceleration-related data from the body of the vehicle during the potential tampering event with a threshold; and in response to determining that the acceleration-related data exceeds the threshold, place the alarm system in a heightened state of alert based on the vehicle tampering event classification.
[0149] According to an embodiment, the present invention is further characterized by instructions for: in response to determining that the acceleration-related data does not exceed the threshold, discarding the vehicle tampering event classification and maintaining the default state of alert of the alarm system.
[0150] According to an embodiment, the present invention is further characterized by instructions for: receiving first audio data from the potential vehicle tampering event from an audio sensor; encoding the first audio data into a latent representation; and partitioning the latent representation into time window frames and stacking the time window frames to generate the time audio data.
[0151] According to an embodiment, the features of the present invention further lie in instructions for performing the following operations: receiving first acceleration-related data from the body of the vehicle during the potential tampering event from an accelerometer; wherein the acceleration-related data is derived from the first acceleration-related data.
[0152] According to an embodiment, the audio sensor is located within the interior space of the vehicle; the accelerometer is mounted to the surface of the body of the vehicle; and the surface of the body of the vehicle interfaces with the interior space of the vehicle.
Claims
1. A method for protecting property, the method comprising: feeding audio data from potential tampering events to a machine learning model trained using audio features of known tampering events; in response to the machine learning model classifying the potential tampering event as a tampering event, comparing acceleration-related data from a subject of the property during the potential tampering event to a threshold; as well as In response to determining that the acceleration-related data exceeds the threshold, an alarm system is placed in a high-alert state based on the tamper event classification.
2. The method of claim 1, further comprising: In response to determining that the acceleration-related data does not exceed the threshold, the tamper event classification is discarded and a default armed state of the alarm system is maintained.
3. The method of claim 1, wherein: The audio data comprises temporal audio data; and The machine learning model includes a temporal convolutional network (TCN) model.
4. The method of claim 3, further comprising: receiving first audio data from the potential tampering event; encoding the first audio data into a latent representation; dividing the latent representation into time window frames and stacking the time window frames to generate the temporal audio data; and The temporal audio data is provided to the TCN model.
5. The method of claim 1, further comprising: receiving first audio data from the potential tampering event from an audio sensor of the alarm system; processing the first audio data to generate the audio data; as well as During the tamper event, the acceleration-related data from the body of the property is received from an accelerometer of the alarm system.
6. The method of claim 5, wherein: said property comprises a vehicle, and said body of said property comprises the body of said vehicle; The audio sensor is located within the interior space of the vehicle; the accelerometer being mounted to a surface of the body of the vehicle; and The surface of the body of the vehicle interfaces with the interior space of the vehicle.
7. The method of claim 1, wherein the tamper event classification comprises at least one of: Classification of handle pull tampering events; Classification of borehole tampering incidents; Key-lock tampering event classification; Classification of metal stripping related tampering events; and Catalytic converter theft tampering incident classification.
8. The method of claim 1 , wherein placing the alarm system in the high alert state comprises at least one of: additional sensors for activating said alarm system; Activate an audio alarm; Activate a visual alarm; and An alarm notification is sent to a location remote from the alarm system.
9. An alarm system comprising one or more processing resources and a memory coupled to the one or more processing resources, the memory storing instructions executable by the one or more processing resources to implement the method of any one of claims 1 to 8.
10. A non-transitory computer-readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to implement the method of any one of claims 1 to 8.