System and method for determining acoustic characteristics of gun and projectile combinations

By using deep neural networks and trainable weighted masks on consumer-grade devices, the problem of rapid identification of acoustic features of gun and ammunition combinations is solved, and accurate gun and ammunition combination recognition is achieved in complex environments, improving the efficiency and accuracy of shooting recognition.

CN120239804APending Publication Date: 2025-07-01PRECISION SHOOTING CO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380063465.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-02
Filing Date
2023-09-04
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

The prior art is difficult to rely solely on consumer-grade audio recording devices to quickly and accurately determine the acoustic characteristics of guns and ammunition combinations, especially in complex environments.

Method used

Deep neural networks (DNNs) are used to combine trainable weighted masks to generate spectrum by receiving audio signals fired by guns, and use masked spectrum inputs to classify them. Consumer-grade devices such as smartphones are used for real-time or near-real-time gun and ammunition combination identification.

Benefits of technology

It realizes rapid and accurate identification of gun and ammunition combinations in complex environments, reduces false alarms and omissions in scoring systems, and is suitable for automatic target scoring in crowded shooting ranges.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120239804A_ABST
    Figure CN120239804A_ABST
Patent Text Reader

Abstract

A system is configured to receive an audio record of firearm firing, calculate a spectrum based on the audio record, and determine a particular firearm and ammunition combination associated with the audio record. The deep neural network may be executed by one or more processors, which may be trained to calculate a spectrum associated with firing audio, derive a spectrum weighting, and determine a firearm and / or ammunition combination associated with the audio recording. A trained input layer may be applied to input data to improve the speed and accuracy of the determination.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross - reference

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 403,696, filed on September 2, 2022, the entire disclosure of which is incorporated herein by reference.

[0003] Background

[0004] The field of the present disclosure relates to determining the acoustic signature of a firearm and projectile combination. In many cases, it may be desirable to identify a particular firearm based on acoustic information. Additionally, it may be further desirable to determine a particular ammunition and firearm combination. Such systems have many use cases, for example, for determining the order of shots fired during a crime, or for identifying weapons on a battlefield, or for identifying individual shooters at a firing range, such as for scoring purposes.

[0005] Historically, there has not been a system capable of identifying a specific firearm and / or ammunition combination, especially one that can rely solely on consumer - grade audio recording devices.

[0006] Accordingly, there is a need for a system and method capable of determining the acoustic signature of a firearm and that can further determine the characteristics of a firearm and ammunition combination. There is also a need for a system that can use consumer - grade audio recording devices and provide the above - mentioned benefits in near real - time. From the following disclosure, these and other benefits will become apparent.

[0007] Overview of Embodiments

[0008] A system having one or more computers can be configured to perform particular operations or actions by installing software, firmware, hardware, or a combination thereof on the system, where in operation the software, firmware, hardware, or a combination thereof causes the system to perform the actions. One or more computer programs can be configured to perform particular operations or actions by including instructions that, when executed by a data - processing device, cause the device to perform the actions. One general aspect includes a method for determining the acoustic signature of a firearm. The method further includes: receiving an audio signal associated with the discharge of a firearm; generating a spectrum associated with the audio signal; applying a mask to the spectrum to generate a masked - spectrum input; performing a deep neural network configured to classify the masked - spectrum input; and determining a firearm and ammunition combination associated with the audio signal based at least in part on the masked - spectrum input. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, all of which are configured to perform the actions of the method.

[0009] Implementations may include one or more of the following features. In the method, an audio signal is received via a microphone associated with a smartphone. The method is performed on the smartphone. The method may include training a mask based on training data associated with a firearm and ammunition. The method may include receiving subsequent audio data associated with a firing of a second firearm and determining that the firing of the second firearm is not associated with the firearm and ammunition combination associated with the audio signal. The method may include receiving a plurality of audio samples associated with a plurality of firearm firings and identifying, from the plurality of audio samples, individual audio samples associated with the firearm and ammunition combination. Implementations of the described techniques may include hardware, a method or process, or computer software on a computer-accessible medium.

[0010] One general aspect includes a system for determining the acoustic signature of a firearm, including a computing device having one or more processors configured with instructions that may include: a spectrum module configured to receive an audio signal and output a spectrum associated with the audio signal; an attribute tagger configured to tag the spectrum associated with the audio signal with an attribute to generate a tagged spectrum; a weighting module configured to aggregate and weight the tagged spectrum to generate a weighted spectrum; and a deep neural network classifier configured to classify the weighted spectrum to determine the firearm associated with the audio signal. Other embodiments of this aspect include corresponding computer systems, apparatus, and computer programs recorded on one or more computer storage devices, each configured to perform the actions of the method.

[0011] Implementations may include one or more of the following features. In the system, the computing device is a smartphone. The system may include a microphone configured to capture an audio signal. The system may include a camera configured to capture an audio / video signal. The weighting module is further configured to train a weighting mask based on the spectrum to generate the acoustic signature of the firearm. The system is configured to distinguish a first acoustic signature associated with a first firearm from a second acoustic signature associated with a second firearm. Implementations of the described techniques may include hardware, a method or process, or computer software on a computer-accessible medium. Brief Description of the Drawings

[0013] The accompanying drawings are part of this disclosure and are incorporated into this specification. The accompanying drawings illustrate examples of embodiments of this disclosure and, in conjunction with the specification and the claims, are used to at least partially explain various principles, features, or aspects of this disclosure. Some embodiments of this disclosure are described more fully below with reference to the accompanying drawings. However, the various aspects of this disclosure may be implemented in many different forms and should not be construed as limited to the implementations set forth herein. Like numerals refer to like but not necessarily identical or the same elements throughout.

[0014] Figure 1 illustrates an example system for determining acoustic features according to some embodiments;

[0015] Figure 2 illustrates an example flowchart for pre-training a DNN according to some embodiments;

[0016] Figure 3 illustrates an example flowchart for deriving spectral weighting according to some embodiments; and

[0017] Figure 4 illustrates an example flowchart for classifying and determining acoustic features.

[0018] Detailed Description

[0019] According to some embodiments, a system is described that can rapidly compute the characteristics of the acoustic sound of a shooter's weapon firing a specific ammunition in a given environment using a consumer-grade audio recording device (e.g., iPhone, tablet, cell phone, video camera). In some cases, the system can receive an input audio or audio / video file and compute the characteristics of the firearm based on the sounds in the captured recording. The recording can be captured by any suitable audio and / or video capture device, such as but not limited to security cameras, traffic cameras, video cameras, television cameras, mobile device recorders such as smartphones or tablets, and other capture devices. In some cases, the capture device is an easily obtainable consumer-grade recording device. The acoustic characteristics can also include determining the brand, model, silencer, ammunition type, ammunition manufacturer, and other characteristics of the weapon explosion based on the acoustic characteristics from the fired firearm.

[0020] This ability can be used to detect a shooter's shot and distinguish it from the shots of other shooters on a typical shooting range, which may be useful, for example, for automatic scoring. One form of this ability can be used to identify a specific weapon and ammunition from a recording.

[0021] In some cases, the described solution operates near real-time on a single consumer recording device such as a mobile phone using only a modest amount of training data. As used herein, the term "real-time" or "near real-time" is a broad term and, in the context of the present disclosure, relates to receiving input data, processing the input data, and outputting the results of the data analysis with little to no perceptible latency to a human. In other words, a system that outputs analyzed data in less than one second as described herein is considered near real-time. A system that operates in real-time or near real-time can limit the computational effort of machine learning or at least a trained model that the method can use to characterize a particular weapon that fires appropriate ammunition. Additionally, there are methods that do not use ML techniques, which means they do not adjust ("train") a model based on large numbers of samples. Some of these other methods can work well, but can also be computationally intensive or shift the computational load to the shot identification time, e.g., such as searching for a "match" in a large feature library.

[0022] Existing methods for analyzing gunshots using artificial intelligence (AI) propose a straightforward approach that uses a deep neural network (DNN) to classify gunshot sounds by weapon and ammunition type. In existing methods, the DNN must be pre-trained based on a large number of gunshot sounds for each weapon and appropriate ammunition type. These sounds may not have been captured on the same recording device or under the same shooting conditions. This potentially increases the generalization ability of the trained DNN instance to classify weapon and ammunition types regardless of the environment and recording device. However, it reduces its ability to identify gunshot sounds of a specific weapon and ammunition type using any recording device in any environment.

[0023] Existing methods describe a two-step approach to classification. For example, some existing methods can use a pre-trained instance of a relatively generic pre-trained DNN instance as an approximate classifier (predictor) and use an additional trainable step to refine the DNN prediction based on the classifier output. In some cases, the DNN treats a finite-length time segment of the time-frequency spectrogram of the gunshot sound as an image and differentiates the pooled images for each weapon and ammunition pair from the pooled images of other weapon and ammunition pairs. This method has several drawbacks, including using a generic DNN instance that is not particularly good at classifying sounds from different environments or different ammunition.

[0024] According to some embodiments, the described system superimposes a trainable weighted mask on an image created from time segments of a time-frequency spectrogram, along with a straightforward training method for adjusting the mask using only a few instances of the shooter's weapon and ammunition firing sounds recorded with a specific device in a specific environment. Compared to training the DNN itself, training only the weighted mask is significantly less computationally intensive. Thus, according to some embodiments, the DNN may not be trained, or may be trained to a much lesser extent than existing methods, and the weighted mask receives training. This approach has several benefits.

[0025] For example, abstractly speaking, the adjusted mask can be regarded as a feature for the combination of weapon, ammunition, environment, and recording device. This feature is not used to search a catalog of features for this combination of factors, but rather to adjust the input spectrum to enhance the DNN's classification performance for the weapon and ammunition in a specific environment using a given recording device. The enhanced DNN classification can then be used to improve the detection of the shooter's shots and distinguish them from the shots of other shooters.

[0026] This approach shows significant advantages in terms of the intensity of the required computations, leading to faster analysis, and can be used for many guns in many different environments. For example, since the training dataset for a specific shooter is small, the weighted mask can be updated online using stochastic gradient descent or offline using gradient descent. In other words, the training data can be specific to a particular shooter using a specific combination of weapon and ammunition in a given environment, resulting in a much smaller dataset than the case of aggregating data points from numerous shooters in different environments.

[0027] Furthermore, both algorithms can estimate the gradient of the function computed by the DNN in each iteration of the weighted matrix update. While this is not computationally prohibitive, it can be found that only the sign of the gradient or a roughly quantized form is sufficient. This simplification depends to a large extent on whether the DNN computes a monotonically increasing or decreasing function for each input. There is a large body of literature studying the properties of the functions computed by DNNs. However, it is not easy to find a short answer to this question. In some cases, the DNN does compute a monotonically increasing or decreasing function for each input if all internal layers use linear or affine activation functions and the output layer uses a monotonically non-decreasing activation function.

[0028] Thus, in some embodiments, the disclosed system is capable of rapidly fingerprinting weapons and ammunition pairs in a particular environment using a particular recording device. In other words, the systems and methods described herein can very quickly determine the acoustic signature of a firearm and ammunition combination in an environment. This allows the system to distinguish the analyzed acoustic signature from other firearm and ammunition combinations. This is particularly useful in a crowded shooting range where it is important to distinguish one shooter from another, e.g., in combination with a system that performs automatic target scoring. By being able to distinguish firearm and ammunition pairs, the number of false positives and missed shots of the scoring system is reduced because the system will be able to determine that a particular weapon and ammunition combination was used, which can be matched in time to a hit on a target. In some examples, the system can be executed on a mobile computing device and used at a shooting range. In the case where the mobile computing device has a microphone pointed in the general direction of the shooter of interest, the shot fired by the shooter of interest will typically have an audio file dominated by the shot fired by the shooter of interest, which can help determine whether the shot fired came from the shooter of interest. In some cases, the recording device (e.g., the mobile computing device) has a microphone pointed at the shooting range, such as in the case where the mobile computing device has a camera pointed at the target (such as for automatic target scoring), in which case the audio of the shot fired by the shooter of interest will have a volume that is more difficult to distinguish from the other shooters at the shooting range. In these cases, the described embodiments can use training of machine learning algorithms or training of weighted masks to quickly distinguish shots fired by the intended shooter from all other shots fired on the range.

[0029] In some cases, the classifier relies on feature extraction in the time - frequency domain. Mel - Frequency Cepstral Coefficient (MFCC) vectors can be used as an alternative to the raw Power Spectral Density (PSD) vector for the frequency representation of the spectrum. Initially, Mel - Frequency Cepstrum (MFC) was developed for speech processing, based on the way information was thought to be encoded in the speech waveform. In some cases, MFC can be applied to firearm fingerprinting purposes. As used throughout this disclosure, "firearm fingerprint" is used to refer to the acoustic signature, or the particular sound or sound image that identifies a particular firearm and ammunition combination.

[0030] In some cases, MFCCs can be generated by performing a series of steps, including but not limited to: i) windowing a signal segment and computing the fast Fourier transform (FFT) of the signal segment, ii) combining the linear FFT coefficients into Mel frequency filter bank coefficients, iii) obtaining the logarithms of these coefficients, and iv) computing the discrete cosine transform (DCT) of the logarithmic Mel filter bank coefficients. The FFT of a signal segment typically results in peaks at the applied frequencies and additional peaks, often referred to as sidelobes, on either side of the peak frequencies. In some cases, the DCT expresses a finite sequence of data points as a sum of cosine functions oscillating at different frequencies. In some cases, fewer or more steps than those disclosed may be implemented to derive a firearm fingerprint based on a gunshot sound. For example, in some cases, for a gunshot sound, only steps i), ii), and iii) above may be used.

[0031] According to some embodiments, MFCC or PSD coefficients can be used as inputs to a DNN, which is trained to classify time-spectra into a number of distinct classes. For example, MFCC and / or PSD coefficients can be input into two classes as suggested, or a number of classes corresponding to (weapon, ammunition) pairs, or a larger number of classes corresponding to tuples with k > 2 attributes.

[0032] In some cases, the DNN can place a new gunshot sound "near" any class represented by the maximum classifier output among all classifier outputs. For example, a gunshot sound can be initially classified by a nearest neighbor method such as the k-NN algorithm. Subsequent analysis can further classify the gunshot sound.

[0033] The inner layers of the DNN can represent different sets of attributes of the input. These can likewise be used to classify and train gunshot sounds.

[0034] From an optimization theory perspective, training a DNN essentially defines a surface with multiple local optima, and the DNN can be considered to guide new inputs to the most appropriate local optimum. In some embodiments, connecting one or more trainable layers after a pre-trained DNN can be considered to extract different sets of more optimal attributes for gunshot sounds.

[0035] In some cases, connecting a trainable layer (which may include weighted coefficients in some cases) before a pre-trained DNN can be considered as adjusting the pre-trained DNN such that the shot of the shooter of interest is the most positive example among all shots in the pre-trained DNN and is placed "near" any classification represented by the maximum classifier output among all classifier outputs. In some cases, the decision threshold can be adjusted to optimize the confusion matrix for the training dataset used to adjust the trainable input layer or for an additional test dataset regarding some useful criteria.

[0036] According to some embodiments, the FFT of the windowed segment of the shot audio can be computed in O(n log n) time. Computing the MFCC has the same time order but includes computing two O(n log n) operations. In some cases, the MFCC may not be significantly better than the FFT and they can be omitted.

[0037] According to some embodiments, the mel-frequency log-spectrum (omitting the discrete cosine transform (DCT) that produces the cepstrum) can be an improvement over the original FFT with only an additional O(n) computational cost. Thus, in some examples, the mel-frequency log-spectrum is used instead of the DCT that produces the cepstrum.

[0038] Generally, there is no fast form like the FFT for computing the function computed by a DNN (composed of convolutional neural network (CNN) layers). The FFT exploits regularities in the FFT kernel that generally do not exist in the composition of an essentially arbitrary CNN. However, the examples described herein have the FFT of the input time signal and a pre-trained DNN composed of CNNs can be implemented in the frequency domain.

[0039] As a non-limiting example, the following DNN-based fast-trainable classifiers can be implemented to quickly determine the acoustic characteristics of firearm and ammunition combinations.

[0040] Let f: R M →R K represent the simulated transformation of a trained deep neural network from a real-valued M-dimensional input vector to a K-dimensional vector of classification probabilities. The final discrete output mapping Ψ: R K →N ≤K selects the most likely K classification.

[0041] The system can be configured to extend a trained DNN into an enhanced binary classifier for class k. In some cases, the system can add an input stage that can be quickly trained to the DNN, implemented as the Hadamard product "⊙" of a weighted matrix W and an input vector x. In some cases, the Hadamard product is a binary operation that takes two matrices (such as the weighted matrix W and the input vector x) and returns a matrix of the corresponding elements that are multiplied. Then, the system can follow the DNN classification probability vector output with a selector function:

[0042] Γ: R K ×N ≤K →R, providing a single classification probability for class k. The enhanced DNN can implement the following function:

[0043] y(n) = Γ(f(W⊙x(n)); k).

[0044] Suppose we have a set of additional input vectors ζ = {Z1, ……, Z M}, all instances of the same class k. The system can be modulated in any of many suitable ways to better identify similar instances of class k. For example, one method is online learning, which is commonly used for large training datasets. For online learning, we initially set W = 1 and then update it sequentially for the vectors in ζ using stochastic gradient descent as follows:

[0045]

[0046] where 0 < α < 1 is an adaptation constant that weights the relative contribution of ζ to W. For each vector in ζ, the target value d(n) = 1.0. The iteration is performed until it converges for each z in ζ. Here, relu[...] represents the element-wise relu() of the parameter matrix.

[0047] Gun firing results in multiple acoustic events, for example, the muzzle blast generated by the expansion of the gas in the chamber and discharged through the barrel, and the ballistic shock wave generated by the projectile, which is supersonic in most cases but can also be subsonic in some cases. The acoustic events are caused by variables that generate gun characteristics and can include gun type, brand, model, barrel length, ammunition type, gunpowder quantity and composition, projectile weight, and projectile shape, etc.

[0048] Reference Figure 1, which shows an example system 100 using online learning, where audio 102 is received and converted to an incremental spectrum 104. The spectrum 104 is framed 106, for example, by a time window, and used to determine MFCCs, for example, by determining the FFT of the windowed signal, combining the linear FFT coefficients into MEL frequency filter bank coefficients, determining the logarithm of the coefficients, and determining the DCT of the logarithmic MEL filter bank coefficients. The determined MFCCs can be input into a quickly trainable input stage 108 and then delivered to a DNN multi-class classifier 110. A selector 134 determines the classifier output k with the highest classification probability for classifying a shot. The quickly trainable input stage 108 can be referred to as a trainable weighted mask or simply as a mask. The mask can be adjusted using examples of the firing sounds of the shooter's weapon and ammunition recorded in a specific environment. In some cases, the computational intensity of training only the weighted mask is significantly less than training the DNN. The trained (e.g., adjusted) weighted mask can represent the characteristics of a combination of the weapon, ammunition, environment, and recording device. This characteristic is not necessarily used to search a feature catalog but is used to adjust the input spectrum to enhance the DNN's classification performance for weapons and ammunition in a specific environment using a given recording device. In some cases, the training dataset for a specific shooter (e.g., a specific combination of firearm and ammunition) is small, so an online method (such as using stochastic gradient descent) or an offline method (such as using gradient descent techniques) can be used to update the weighted mask.

[0049] The inner layers of the DNN classifier 110 can represent different sets of attributes of the input. In some cases, the DNN is trained to define a surface with multiple local optima, and the DNN can be used to guide new inputs to the most appropriate local optimum. Then, enhanced DNN classification can be used to improve the detection of shots from one weapon and distinguish them from shots from other shooters.

[0050] The weighted mask W of the input stage 108 to the pre-trained DNN classifier 110 can be adjusted through an online learning loop 112, and the online learning loop 112 optimizes the weighted mask W.

[0051] The online learning loop 112 includes copies 114, 116, and 118 of the shot classifiers 108, 110, and 134. When a new shot n is detected, the subtractor 120 in the online learning loop 112 compares the result of the selected output k from the copies 114, 116, and 118 of the shot classifier with the expected classification d(n) = 1.0 to calculate the classification error term.

[0052] The block 130 of the adjustment loop calculates the vector sign of the gradient of the DNN output with respect to the spectral input of the current output spectrum x(n) from the framing block 106.

[0053] For the current weighted vector W i (n), the original incremental adjustment value is calculated by the Hadamard multiplier 132 as the product of the current firing spectrum from the frame divider 106 and the sign vector of the classifier gradient 130. Then, the vector multiplier 122 scales the original incremental adjustment value by multiplying the weighted vector W i (n) by an arbitrary value α.

[0054] Finally, the vector adder 124 calculates the preliminarily updated weight by adding the current incremental adjustment value to the current W i (n). This preliminary weight is transformed by the vector relu[] operation 126 into the updated weighted vector W i+1 (n).

[0055] The weight adjustment loop 112 just described is iterated, as indicated by the Δ operator 128, until the weighted vector converges, e.g., |W i (n) - W i+1 (n)| ≤ δ. Then, the resulting W i+1 (n) is selected as the weighted vector W(n + 1) for classifying the next firing.

[0056] Offline learning can be used, e.g., when ζ is very small. In some cases, offline learning can be initiated by setting W = 1 and updating it using gradient descent:

[0057]

[0058] In some cases, offline learning may require approximately the same computational effort as online learning. However, online learning has an advantage over offline learning: Stochastic gradient descent does not require accessing the entire training data set ζ in every iteration of the update function, while offline learning does. Offline learning offers the advantage of finding local optima, while online learning may only approximate it.

[0059] Although the computational cost of evaluating the gradient l ⊙z i at the current parameter W is not overly high, it may be sufficient to pre-compute or where 1 is the unit vector if the DNN normalizes the input. In some cases, this can significantly simplify offline learning in this case, at the cost of only approximating the local optimum as in online learning.

[0060] While there are some systems designed to enhance DNNs for classifying gunshots, they are customized for specific tasks by adding layers to the output of a pre-trained network. The pre-trained network extracts increasingly abstract features at many levels, such as from an image, and additional trainable layers are used to focus on the issues of interest. In contrast, many of the systems and methods described in this article work in a very different way, resulting in more efficient, faster, and more accurate systems. In many cases, the systems described in this article add only a single layer to the input of the pre-trained network. In some use cases, the camera is pointed at the target rather than focused on the shooter, and the camera and microphone pick up other gunshots and cannot natively determine that the firing originated from the shooter of interest. Acoustic features are generated quickly and, in many cases, are performed on a mobile device (e.g., a smartphone) that includes a camera and a microphone. The mobile device can execute instructions (e.g., an application) that include the components and systems described in this article, such that classification and acoustic feature determination are performed on the mobile device. In some cases, the pre-trained network can be trained based on a relatively comprehensive gunshot dataset. The trainable input layer can pre-distort the input data to enable the pre-trained network to highly probably identify a specific weapon as well as ammunition and the environment. This may then increase the likelihood of detecting the gunshot of interest and eliminating all other gunshots.

[0061] According to some embodiments, the systems and methods described in this article will converge to a useful weighted mask where the gradient of the function computed by the DNN is monotonically non-decreasing or monotonically non-increasing in each input.

[0062] Refer to Figure 2 , which shows a pre-trained DNN. Process 200 starts at 202 and opens an audio file at block 204, which can be the first audio file, or the next or subsequent audio file. Using the audio file, and at step 206, the system captures a sample block that frames the first gunshot and / or the next gunshot. In other words, each gunshot is confined to samples within a time range. At block 208, a spectrum associated with the samples is generated and labeled with attributes.

[0063] At block 210, the system determines whether the most recent sample is associated with the last gunshot, and if not, the system returns to block 206 to capture a sample block associated with a subsequent gunshot. If so, the system advances to block 212 and determines whether the labeled spectrum created at block 208 is the last file. If not, the system returns to block 204 to open or capture the next audio file. If the system determines that the most recent file is the last file, the system advances to block 214, where the labeled spectra are aggregated. At block 216, the DNN is trained based on the labeled spectra. At block 218, the system stops and has a trained DNN.

[0064] Reference Figure 3 , which shows the derivation of spectral weighting. Process 300 starts at 302 and captures a set of shots at block 304. The shots can be captured by an audio and / or video recording device or can include opening a file associated with one or more shots. At block 306, the system captures a block of samples that frames the first shot and / or the next shot. In other words, each shot is confined to samples within a time range. At block 308, a spectrum associated with the samples is created and labeled with attributes.

[0065] At block 310, the system determines whether the most recent spectrum is associated with the last shot, and if not, the system returns to block 306 to capture a block of samples associated with a subsequent shot. If so, the system advances to block 312 and determines whether the labeled spectrum created at block 308 is the last set. If not, the system returns to block 304 to capture the next set of shots. If the system determines that the most recent file is the last file, the system advances to block 314, where the labeled spectra are aggregated. At block 316, the system updates W (weighting) until the value converges. At block 218, the system stops and has spectral weighting.

[0066] Reference Figure 4 , which shows the process for classifying shot 400. The process starts at block 402 and at block 404, the system captures a block of samples that frames the shot. At block 406, the system determines the spectrum associated with the block of samples. At block 408, the system determines the Hadamard product of the spectrum and the weight matrix. At block 410, the Hadamard product of the spectrum and the weight matrix is applied to the DNN. At block 412, the system makes a binary decision, such as whether the shot is associated with the characteristics of the firearm in question or not. In some cases, the system is able to identify the type of firearm and the ammunition fired through the firearm. For example, when an audio sample is received, the system can determine that the acoustic characteristics of the firearm in the audio sample correspond to a 230 grain round nose bullet fired from a Beretta.45 ACP without any prior knowledge of the firearm in question. At block 414, the process stops.

[0067] The system can include one or more processors and one or more computer-readable media that can store various modules, applications, programs, or other data. The computer-readable media can include instructions that, when executed by one or more processors, cause the processors to perform the operations described herein for the system.

[0068] In some implementations, the processor may include a central processing unit (CPU), a graphics processing unit (GPU), both a CPU and a GPU, a microprocessor, a digital signal processor, or other processing units or components known in the art. Alternatively or additionally, what is functionally described herein may be performed at least in part by one or more hardware logic components. For example, but not limited to, illustrative types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), etc. Additionally, each of the processors may have its own local memory, which may also store program modules, program data, and / or one or more operating systems. One or more control systems, computer controllers, and remote controls may include one or more cores.

[0069] Embodiments may be provided as a computer program product including a non-transitory machine-readable storage medium having instructions (in compressed or uncompressed form) stored thereon that may be used to program a computer (or other electronic device) to perform the processes or methods described herein. Computer-readable media may include volatile and / or non-volatile memory, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Machine-readable storage media may include, but are not limited to, hard disk drives, floppy disks, optical disks, CD-ROMs, DVDs, read-only memory (ROM), random access memory (RAM), EPROMs, EEPROMs, flash memory, magnetic or optical cards, solid state memory devices, or other types of media / machine-readable media suitable for storing electronic instructions. Additionally, embodiments may be provided as a computer program product including a transient machine-readable signal (in compressed or uncompressed form). Examples of machine-readable signals, whether or not modulated by a carrier, include, but are not limited to, signals that a computer system or machine hosting or running a computer program may be configured to access, including signals downloaded via the Internet or other networks. In some examples, the process is performed on a mobile electronic device such as a smartphone, tablet computer, laptop computer, or other mobile computing device. For example, a smartphone may download and run an application (app) that configures the smartphone to capture sounds and / or audio-video associated with a firearm using a built-in microphone and / or camera, and then perform the methods described herein to classify a shot and determine whether the shot is associated with a particular firearm and ammunition combination.

[0070] Those of ordinary skill in the art will recognize that any process or method disclosed herein can be modified in a variety of ways. The process parameters and the order of steps described and / or illustrated herein are given by way of example only and can be varied as needed. For example, although the steps shown and / or described herein may be shown or discussed in a particular order, these steps do not necessarily need to be performed in the order shown or discussed.

[0071] The various exemplary methods described and / or illustrated herein may also omit one or more steps described or illustrated herein, or include additional steps in addition to those disclosed steps. Furthermore, the steps of any method disclosed herein can be combined with any one or more steps of any other method disclosed herein.

[0072] This disclosure sets forth exemplary embodiments, and thus is not intended to limit in any way the scope of the embodiments of this disclosure and the appended claims. The embodiments have been described above with the aid of functional building blocks that illustrate the implementation of specific functions and their relationships. For ease of description, the boundaries of these functional building blocks have been arbitrarily defined herein. Alternative boundaries can be defined to the extent that appropriately performs the specified functions and their relationships.

[0073] The foregoing description of specific embodiments will so fully disclose the general nature of the embodiments of this disclosure that others can, by applying the knowledge of those of ordinary skill in the art, readily modify and / or adapt various applications of such specific embodiments without undue experimentation, without departing from the general concept of the embodiments of this disclosure. Therefore, such adaptations and modifications are intended to be within the meaning and scope of the equivalents of the disclosed embodiments based on the teachings and guidance presented herein. The words or terms of this specification are for the purpose of description and not of limitation, such that the terms or words of this specification will be interpreted by those of ordinary skill in the relevant art in accordance with the teachings and guidance presented herein.

[0074] The breadth and scope of the embodiments of this disclosure should not be limited by any of the exemplary embodiments described above, but should be defined only in accordance with the appended claims and their equivalents.

[0075] Conditional language, such as “can,” “could,” “might,” or “may,” and others, unless specifically stated otherwise or otherwise understood in the context in which it is used, generally intends to convey that certain embodiments may include certain features, elements, and / or operations, while other embodiments do not include certain features, elements, and / or operations. Thus, such conditional language generally does not intend to imply that the features, elements, and / or operations are required in any way for one or more embodiments, or that one or more embodiments must include logic for determining whether these features, elements, and / or operations are included in or will be performed in any particular embodiment with or without user input or prompting.

[0076] Unless otherwise specified, the terms “connected to” and “coupled to” (and their derivatives) as used in the specification should be interpreted to allow both direct connection and indirect connection (i.e., via other elements or components). Additionally, the term “a” or “an” as used in the specification should be interpreted to mean “at least one of...” Finally, for ease of use, the terms “comprising” and “having” (and their derivatives) as used in the specification do not exclude additional components and should be interpreted as open-ended.

[0077] The specification and drawings disclose examples of systems, devices, apparatuses, and techniques that can provide systems and methods for determining the acoustic signature of a fired firearm. Of course, it is not possible to describe every conceivable combination of elements and / or methods for the purpose of describing the many features of the present disclosure, but one of ordinary skill in the art will recognize that many additional combinations and permutations of the disclosed features are possible. Thus, various modifications may be made to the present disclosure without departing from the scope or spirit thereof. Additionally, other embodiments of the present disclosure may be apparent by considering the specification and drawings, as well as the practice of the disclosed embodiments as presented herein. The examples set forth in the specification and drawings should be considered illustrative in all respects and not restrictive. Although specific terms are employed herein, these terms are used in a general and descriptive sense only and not for purposes of limitation.

[0078] Those skilled in the art will understand that in some embodiments, the functionality provided by the above-described processes and systems can be provided in alternative ways, such as being split among more software programs or routines, or combined into fewer programs or routines. Similarly, in some implementations, the illustrated processes and systems can provide more or less functionality than described, such as when other illustrated processes alternatively lack or include such functionality respectively, or when the amount of functionality provided is changed. Additionally, although various operations may be shown as being performed in a particular manner (e.g., serially or in parallel) and / or in a particular order, those skilled in the art will understand that in other implementations, the operations can be performed in other orders and in other ways. Those skilled in the art will also understand that the data structures discussed above can be constructed in different ways, such as by splitting a single data structure into multiple data structures or by combining multiple data structures into a single data structure. Similarly, in some implementations, the illustrated data structures can store more or less information than described, such as when other illustrated data structures alternatively lack or include such information respectively, or when the amount or type of information stored is changed. The various methods and systems shown in the drawings and described herein represent example implementations. In other embodiments, the methods and systems can be implemented in software, hardware, or a combination thereof. Similarly, in other implementations, the order of any method can be changed, and various elements can be added, reordered, combined, omitted, modified, etc.

[0079] In view of the foregoing, it should be understood that although specific implementations have been described herein for purposes of illustration, various modifications can be made without departing from the spirit and scope of the appended claims and the elements recited therein. Additionally, although certain aspects are presented above in the form of certain claims, the inventors contemplate multiple aspects in any available claim form. For example, although only some aspects may currently be recited as being embodied in a particular configuration, other aspects can equally be so embodied. As will be apparent to those skilled in the art who benefit from this disclosure, various modifications and changes can be made. It is intended to embrace all such modifications and changes, and thus, the above description should be regarded as illustrative rather than restrictive.

Claims

1. A method for determining the acoustic signature of a firearm, comprising: Receiving an audio signal associated with the firing of a firearm; Generating a spectrum associated with the audio signal; Applying a mask to the spectrum to generate a masked spectrum input; Executing a deep neural network configured to generate a classification of the masked spectrum input; and Determining a firearm and ammunition combination associated with the audio signal, at least in part based on the masked spectrum input.

2. The method according to claim 1, wherein, Receiving the audio signal is performed by a microphone associated with a smartphone.

3. The method according to claim 2, wherein The method is executed on the smartphone.

4. The method according to claim 1, further comprising training the mask based on training data associated with the firearm and ammunition.

5. The method according to claim 4, wherein The mask is a Hadamard product of a weight matrix and an input vector.

6. The method according to claim 1, further comprising receiving subsequent audio data associated with the firing of a second firearm and determining that the firing of the second firearm is not associated with the firearm and ammunition combination associated with the audio signal.

7. The method according to claim 1, further comprising receiving a plurality of audio samples associated with a plurality of firearm firings and identifying individual audio samples among the plurality of audio samples that are associated with the firearm and ammunition combination.

8. A system for determining the acoustic signature of a firearm, comprising: A computing device having one or more processors configured with instructions including: A spectrum module configured to receive an audio signal and output a spectrum associated with the audio signal; An attribute tagger configured to tag the spectrum associated with the audio signal with an attribute to generate a tagged spectrum; A weighting module configured to aggregate and weight the tagged spectrum to generate a weighted spectrum; and A deep neural network classifier configured to classify the weighted spectrum to determine a firearm associated with the audio signal.

9. The system according to claim 8, wherein The computing device is a smartphone.

10. The system according to claim 8, further comprising a microphone configured to capture the audio signal.

11. The system according to claim 8, further comprising a camera configured to capture audio / video signals.

12. The system according to claim 8, wherein, The weighting module is further configured to train a weighting mask based on the spectrum to generate the acoustic signature of the firearm.

13. The system according to claim 12, wherein, The weighting mask is generated by determining the Hadamard product of a weight matrix and an input vector to generate the weighted spectrum.

14. The system according to claim 8, wherein, The system is configured to distinguish a first acoustic signature associated with a first firearm from a second acoustic signature associated with a second firearm.