Systems and methods for determining the acoustic signature of a firearm and projectile combination

A DNN with a trainable weight mask on consumer devices quickly identifies firearm and ammunition combinations, addressing the challenge of environmental variability and enhancing classification accuracy for shooter identification in shooting ranges.

JP2025531768APending Publication Date: 2025-09-25ACCUSHOOT INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2025513473
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-09-02
Filing Date
2023-09-04
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Existing systems lack the ability to accurately identify specific firearm and ammunition combinations using consumer audio recording devices, particularly in near real-time environments, and existing AI methods are not adept at classifying sounds from different environments or recording devices.

Method used

A system utilizing a deep neural network (DNN) with a trainable weight mask to process audio signals from consumer devices, allowing for rapid identification of firearm and ammunition combinations by conditioning input spectrograms to enhance classification performance in specific environments.

Benefits of technology

The system efficiently distinguishes between firearm and ammunition pairs in various environments, reducing computational intensity and enabling fast, accurate identification of shooter's shots, useful in crowded shooting ranges for automated scoring.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025531768000001_ABST
    Figure 2025531768000001_ABST
Patent Text Reader

Abstract

A system configured to receive an audio recording of a firearm discharge, calculate a spectrogram based on the audio recording, and determine a particular firearm and ammunition combination associated with the audio recording. A deep neural network may be executed by one or more processors that can be trained to calculate a spectrogram associated with the firearm discharge audio, derive spectrogram weightings, and determine the firearm and / or ammunition combination associated with the audio recording. A trained input layer can be applied to the input data to improve the speed and accuracy of the determination.
Need to check novelty before this filing date? Find Prior Art

Description

Detailed Description of the Invention

[0001] [Cross reference] This application claims the benefit of U.S. Provisional Application No. 63 / 403,696, filed September 2, 2022, the entire disclosure of which is incorporated herein by reference. 〔background〕

[0002] The field of the disclosure relates to determining the acoustic signature of a firearm and projectile combination. In many cases, it may be desirable to identify a particular firearm based on acoustic information. Additionally, it may be desirable to identify a particular ammunition and firearm combination. Such a system has many use cases, for example, determining the sequence of shots fired during the commission of a crime, identifying weapons on the battlefield, or identifying different shooters at a shooting range, such as for scoring purposes.

[0003] Historically, there has been no system capable of identifying specific firearm and / or ammunition combinations, especially those that rely solely on civilian audio recording devices.

[0004] Therefore, there is a need for a system and method that can determine the acoustic signature of a firearm, and further that can determine the signature of a firearm and ammunition combination. Further, there is a need for a system that can provide the aforementioned advantages in near real time using consumer audio recording equipment. These and other advantages will be readily apparent from the following disclosure. [Outline of the embodiment]

[0005] One or more computer systems can be configured to perform specific operations or actions by installing software, firmware, hardware, or a combination thereof that, when operated, causes the system to perform or behave in accordance with the operations. One or more computer programs can be configured to perform specific operations or actions by including instructions that, when executed by a data processing device, cause the device to perform the operations. One general aspect includes a method for determining an acoustic signature of a firearm. The method also includes receiving an audio signal associated with the discharge of a firearm, generating a spectrogram associated with the audio signal, applying a mask to the spectrogram to generate a masked spectrogram input, executing a deep neural network configured to generate a classification of the masked spectrogram input, and determining a firearm and ammunition combination associated with the audio signal based at least in part on the masked spectrogram input. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the operations of the method.

[0006] Implementations may include one or more of the following features. A method, wherein receiving an audio signal is performed by a microphone associated with a smartphone. The method is executed on the smartphone. The method may include training the mask with training data associated with the firearm and the ammunition. The method may include receiving subsequent audio data associated with a firearm firing from a second firearm and determining that the firearm firing from the second firearm is not associated with the firearm and ammunition combination associated with the audio signal. The method may include receiving a plurality of audio samples associated with a plurality of firearm firings and identifying, from the plurality of audio samples, individual ones of the plurality of audio samples associated with the firearm and ammunition combination. Implementations of the described techniques may include hardware, methods or processes, or computer software on a computer-accessible medium.

[0007] One general aspect is a system for determining an acoustic signature of a firearm, the system comprising a computing device having one or more processors, the one or more processors configured with instructions, the computing device including: a spectrogram module configured to receive an audio signal and output a spectrogram associated with the audio signal; an attribute labeler configured to label the spectrogram associated with the audio signal with attributes to generate a labeled spectrogram; a weighting module configured to aggregate and weight the labeled spectrograms to generate a weighted spectrogram; and a deep neural network classifier configured to classify the weighted spectrogram to determine a firearm associated with the audio signal. Other embodiments of this aspect include corresponding computer systems, devices, and computer programs recorded on one or more computer storage devices, each configured to perform the operations of the method.

[0008] Implementations may include one or more of the following features: The system wherein the computing device is a smartphone. The system may include a microphone configured to capture an audio signal. The system may include a camera configured to capture an audio / video signal. The weighting module is further configured to learn a weight mask on the spectrogram to generate an acoustic signature of the firearm. The system is configured to distinguish between a first acoustic signature associated with a first firearm and a second acoustic signature associated with a second firearm. Implementations of the described techniques may include hardware, methods or processes, or computer software on a computer-accessible medium. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] The accompanying drawings are part of this disclosure and are incorporated herein. The drawings illustrate example embodiments of the present disclosure and, together with the specification and claims, serve to explain, at least in part, various principles, features, or aspects of the present disclosure. Specific embodiments of the present disclosure are described more fully below with reference to the accompanying drawings. However, various aspects of the present disclosure may be embodied in many different forms and should not be construed as limited to the implementations set forth herein. Like numbers refer to similar, but not necessarily identical, elements throughout.

[0010] FIG. 1 illustrates a sample system for determining an acoustic signature, according to some embodiments.

[0011] FIG. 2 illustrates a sample flowchart for pre-training a DNN, according to some embodiments.

[0012] FIG. 3 illustrates a sample flowchart for deriving spectrogram weightings, according to some embodiments.

[0013] FIG. 4 illustrates a sample flow chart for classifying and determining acoustic signatures. Detailed Description

[0014] According to some embodiments, a system is described that can rapidly calculate the acoustic sound signature of a shooter's weapon firing a specific ammunition in a given environment using a consumer audio recording device (e.g., iPhone, tablet, phone, video camera). In some cases, the system can receive an input audio or audio / video file and calculate a firearm signature based on the audio in the captured recording. The recording may be captured by any suitable audio and / or video capture device, including, but not limited to, security cameras, traffic cameras, video cameras, television cameras, mobile device recording devices such as smartphones or tablets, and other capture devices. In some cases, the capture device is a readily available consumer audio recording device. The acoustic signature may further include determining the make, model, silencer, ammunition type, ammunition manufacturer, and other characteristics of the weapon's blast wave based on the acoustic signature from the fired firearm.

[0015] This feature may be used to detect and separate a shooter's shots from those of other shooters in a typical shooting range, thereby aiding in automatic scoring, etc. A version of this feature may be used to identify a particular weapon or ammunition from a recording.

[0016] In some cases, the described solutions operate in near real time on a single consumer recording device, such as a mobile phone, using only a modest amount of training data. As used herein, the terms “real time” or “near real time” are broad terms and, in the context of this disclosure, refer to receiving input data, processing the input data, and outputting data analysis results with latencies barely perceptible to humans. In other words, a system such as that described herein that outputs analysis data in less than one second is considered near real time. Systems that operate in real time or near real time may limit the computational complexity of machine learning, or at least the training model, that a method can use to characterize a particular weapon firing the appropriate ammunition. Additionally, some methods do not use ML techniques, meaning they do not adapt (“train”) models to a large number of samples. Some of these other methods may work well but may be more computationally intensive or shift the computational burden to shot identification time, for example, by searching a large dictionary of signatures for “matches.”

[0017] Previous methods using artificial intelligence (AI) to analyze gun sounds have introduced a simple method for classifying shots by weapon and ammunition type using deep neural networks (DNNs). In these methods, the DNN must be pre-trained with numerous firearm sounds for each weapon and the appropriate ammunition type. These sounds may not necessarily be acquired using the same recording device or under the same shooting conditions. This potentially improves the generalization ability of the trained DNN instance to classify weapon and ammunition types regardless of the environment and recording device. However, this reduces the ability to identify specific weapon fire sounds and ammunition types in any environment using any recording device.

[0018] Previous methods describe a two-stage approach to classification. For example, some previous methods use a relatively generic pre-trained DNN instance as an approximate classifier (predictor) and can apply an additional trainable step to the classifier output to improve the DNN prediction. In some cases, the DNN treats finite-length time segments of the time-varying spectrogram of gunfire sounds as images and distinguishes pooled images of each weapon and ammunition pair from pooled images of other weapon and ammunition pairs. This approach has several drawbacks, including generic DNN instances. Because of their use, generic DNN instances are not particularly adept at classifying sounds from different environments or different ammunition.

[0019] According to some embodiments, the described system provides a simple training method for overlaying a trainable weight mask onto an image created from a time segment of a time-varying spectrogram and adjusting the mask using a small number of instances of a shooter's weapon and ammunition firing, recorded in a specific environment using a specific device. Training only the weight mask is significantly less computationally intensive than training the DNN itself. As a result, according to some embodiments, the DNN is not trained, or is trained to a much lesser extent than traditional methods, and the weight mask is trained. This method has several advantages.

[0020] For example, in abstract terms, the mask being trained can be thought of as a signature for a combination of weapon, ammunition, environment, and recording device. This signature is not used to search a catalog of signatures for combinations of factors, but rather conditions the input spectrogram to enhance the DNN's classification performance for weapons and ammunition in a specific environment using a given recording device. The enhanced DNN classification can be used to improve the detection and differentiation of a shooter's shots from those of other shooters.

[0021] This method offers significant advantages in the computational intensity required, results in much faster analysis, and can be used for a large number of firearms in many different environments. For example, because the training data set for a particular shooter is small, the weight mask can be updated either online using stochastic gradient descent or offline using gradient descent. In other words, the training data can be for a specific shooter using a specific weapon and ammunition combination in a given environment. This results in a much smaller data set than if data points were aggregated from a large number of shooters in different environments.

[0022] Furthermore, both algorithms may estimate the gradient of the function computed by the DNN at each iteration of the weight matrix update. While this is not computationally impossible, the sign of the gradient, or a roughly quantized version of it, may prove sufficient. This simplification depends in large part on whether the DNN computes a monotonically increasing or monotonically decreasing function at each input. There is a large literature studying the properties of functions computed by DNNs. However, a concise answer to this question is not easy to find. In some instances, this occurs when all of the internal layers of a DNN use linear or affine activation functions, and the output layer uses a monotonically non-decreasing activation function.

[0023] Thus, in some embodiments, the disclosed system can quickly fingerprint weapon and ammunition pairs in a particular environment using a specific recording device. In other words, the systems and methods described herein can very quickly determine the acoustic signature of a firearm and ammunition combination in an environment. This allows the system to distinguish the analyzed acoustic signature from other firearm and ammunition combinations. This is particularly useful in crowded shooting ranges where distinguishing one shooter from another is important, such as in combination with a system that performs automated target scoring. The ability to distinguish between firearm and ammunition pairs allows the scoring system to determine that a specific weapon and ammunition combination is being used, thereby reducing the number of false positives and missed shots, which may be time-coincided with a hit on the target. In some examples, the system can be run on a mobile computing device and used at a shooting range. If the mobile computing device has a microphone pointed in the general direction of the shooter of interest, shots fired by the shooter of interest will generally have a dominant audio profile. This may help determine whether a shot is from the shooter of interest. In some cases, the recording device (e.g., a mobile computing device) has a microphone pointed down the range, such as when the mobile computing device has a camera pointed at the target, as in automated target scoring. In such cases, the described embodiments can use training machine learning algorithms or training weight masks to quickly distinguish shots fired by an intended shooter from all other shooters on the range.

[0024] In some cases, classifiers rely on feature extraction in the form of time-frequency spectrograms. Mel-frequency cepstrum coefficient (MFCC) vectors can be used as an alternative to raw power spectral density vectors (PSDs) for the frequency representation of a spectrogram. Mel-frequency cepstrum (MFCC) was originally developed for audio processing and was considered the way information is encoded in audio waveforms. In some cases, MFCs can also be applied to firearm fingerprinting. As used throughout this disclosure, "firearm fingerprint" refers to an acoustic signature that identifies a particular firearm and ammunition combination, or a specific sound or sound image.

[0025] In this case, MFCCs are generated by performing a series of steps, such as: i) windowing a signal segment, computing a fast Fourier transform (FFT) of the signal segment, ii) combining the linear FFT coefficients with MEL frequency filter bank coefficients, iii) taking the logarithm of those coefficients, and iv) computing a discrete cosine transform (DCT) of the logarithmic MEL filter bank coefficients. The FFT of the signal segment generally produces a peak at the applied frequency, along with other peaks called sidelobes. Sidelobes are typically on one side of the peak frequency. The DCT optionally represents a finite sequence of data points as a sum of cosine functions oscillating at different frequencies. In some cases, fewer or more steps than those disclosed may be implemented to arrive at a firearm fingerprint based on gunshot sounds. For example, in some cases, only steps i), ii), and iii) above may be used for gunshot sounds.

[0026] According to some embodiments, the MFCC or PSD coefficients can be used as input to a DNN trained to classify the time-frequency spectrogram into several independent categories. For example, the MFCC and / or PSD coefficients can input two categories as in the proposal, or several categories corresponding to (weapon, ammunition) pairs, or several larger number of categories corresponding to k>2 tuples of feature values.

[0027] In some cases, the DNN may place a new shot sound in the "neighborhood" of whatever class is represented by the largest classifier output among all classifier outputs. For example, a shot sound may be initially classified by a nearest neighbor method such as a k-NN algorithm. Subsequent analysis may further classify the shot sound.

[0028] The inner layers of the DNN may represent different sets of feature values ​​for the input, which may also be used for classification and training of the projectile sounds.

[0029] From an optimization theory perspective, training a DNN can be thought of as essentially defining a surface with multiple local optima, and the DNN directing new inputs to the most appropriate local optima. In some embodiments, the following pre-trained DNN with one or more trainable layers can be thought of as extracting a more distinct set of optimal attributes of the emitted sound.

[0030] In some cases, prioritizing the pre-trained DNN with a trainable layer may include weighting factors. Tuning the pre-trained DNN may be thought of as placing the shooter's shot of interest "near" whatever class the maximum classifier output of all classifier outputs represents, so that the shooter's shot is the most positive example of all the shots in the pre-trained DNN. In some cases, the decision threshold may be adjusted to optimize the confusion matrix of the training dataset used to train the trainable input layer, or a separate test dataset for some useful criterion.

[0031] According to some embodiments, the FFT of a windowed segment of emitting audio can be computed in O(n log n) time. Computing MFCCs has the same time order but involves two O(n log n) operations. In some cases, MFCCs may not be significantly better than the FFT and may be omitted.

[0032] According to some embodiments, the MEL frequency log spectrum (omitting the discrete cosine transform (DCT) that yields the cepstrum) may provide an improvement over the raw FFT at only an extra computational cost of O(n). Therefore, in some examples, the MEL frequency log spectrum is used rather than the DCT that produces the cepstrum.

[0033] In general, there is no fast form for computing the functions computed by DNNs (composed of layers of convolutional neural networks (CNNs)), such as the FFT. The FFT exploits the regularity of the FFT kernel, which is generally not inherently present in any CNN configuration. However, in the example described herein, a pre-trained DNN composed of a CNN with an FFT of the input time signal can be implemented in the frequency domain.

[0034] As a non-limiting example, the following DNN-based, quickly trainable category recognizer may be implemented to quickly determine the acoustic signature of a firearm and ammunition combination.

[0035]

number

number

[0036] The system may be configured to extend the trained DNN into an enhanced binary classifier for class k. In some cases, the system may perform a Hadamard product of a weight matrix W and an input vector x.

number

[0037]

number

[0038]

number

[0039] a set of additional input vectors, all instances of the same class k

number

number

[0040] where 0<α<1 is an adaptive constant that weights the relative contribution of ζ to W. For all vectors of ζ, a target value (n)=1.0 is used. Iterations are performed for each z in ζ until convergence occurs. Here, relu [...] denotes element-wise relu() of the argument matrix.

[0041] The firing of a firearm results in multiple acoustic events, such as a muzzle blast caused by the expanding gases in the chamber and escaping the barrel, and a ballistic shock wave generated by the projectile. The ballistic shock wave is often supersonic, but in some cases may be subsonic. The acoustic events are the result of variables that create the firearm's signature, which may include the type, make, model, barrel length, ammunition type, powder amount and units, projectile weight, projectile shape, etc.

[0042] Referring to FIG. 1, an exemplary system 100 using online learning is described in which audio 102 is received and converted into an incremental spectrogram 104. The spectrogram 104 is framed 106, such as by a time window, and used to determine MFCCs by determining an FFT of the windowed signal, combining the linear FFT coefficients with MEL frequency filter bank coefficients, determining the logarithm of the coefficients, determining the DCT of the logarithmic MEL filter bank coefficients, and so on. The determined MFCCs are input to a rapidly trainable input stage 108 and transmitted to a DNN multi-class classifier 110. A selector 134 determines the classifier output k with the highest class probability to classify the shot. The rapidly trainable input stage 108 may be referred to as a trainable weight mask, or simply a mask. The mask can be tuned using instances of the shooter's weapon and ammunition firing sounds recorded in a particular environment. In some cases, training only the weight mask can be significantly less computationally intensive than training a DNN. The weight mask that is trained (e.g., tuned) may represent a signature of a combination of weapon, ammunition, environment, and recording device. This signature is not necessarily used to search a catalog of signatures, but rather is used to condition input spectrograms to improve the DNN's classification performance for weapons and ammunition in a particular environment using a given recording device. In some cases, due to a small training dataset for a particular shooter (e.g., a particular firearm and ammunition combination), the weight mask can be updated using either online methods, such as stochastic gradient descent, or offline methods, such as gradient descent.

[0043] The inner layers of the DNN classifier 110 can represent different sets of attributes of the input. In some cases, the DNN is trained to define a surface with multiple local optima, and the DNN can operate to direct new inputs to the most appropriate local optima. Enhanced DNN classification can be used to improve detection of shots from one weapon and to improve discrimination from shots from other shooters.

[0044] The weight mask W of the input stage 108 to the pre-trained DNN classifier 110 may be adapted by an online training loop 112 that optimizes the weight mask W.

[0045] Online training loop 112 includes copies 114, 116, and 118 of shot classifiers 108, 110, and 134. When a new shot n is detected, subtractor 120 in online training loop 112 compares the result of the selected output k from the copies of shot classifiers 114, 116, and 118 with the expected classification d(n)=1.0 to calculate a classification error term.

[0046] Block 130 of the adaptation loop calculates, for the current output spectrogram x(n) from framing block 106, the vector sign of the gradient of the DNN output with respect to the spectrogram input.

[0047] The current weight vector W i The raw incremental adjustment to (n) is then calculated by Hadamard multiplier 132 as the product of the current shot spectrogram from framer 106 and the sign vector of classifier gradient 130. Vector multiplier 122 then multiplies the weight vector W i Scale the raw incremental adjustment to (n) by an arbitrary value α.

[0048] Finally, vector adder 124 adds the current incremental adjustment value to the current W i (n) to calculate the preliminary updated weights. The preliminary weights are added to the updated weight vector W by the vector relu[] operation 126. (i+1) (n)

[0049] The weight adaptation loop 112 just described continues until the weight vector converges, as symbolized by the delta operator 128, e.g.

number

[0050] Offline training can be used, such as when ζ is small. In some cases, offline training can begin by setting W=I and updating using gradient descent.

[0051]

number

[0052] In some cases, offline training requires similar computational effort as online training. However, online training has the advantage over offline training that stochastic gradient descent does not need to access the entire training dataset ζ for each iteration of the update function. Offline training offers the advantage of finding local optima, whereas online training can only find approximations to local optima.

[0053] The current argument

number

[0054] While systems exist that aim to enhance DNNs for shot classification, they do so by adding layers to the output of a pre-trained network to customize it for a specific task. The pre-trained network extracts increasingly more levels of abstraction features from, for example, an image, and additional learnable layers are used to focus on the problem of interest. In contrast, many of the systems and methods described herein work in a much different way, resulting in a more efficient, faster, and more accurate system. The systems described herein often only add a single layer to the input of a pre-trained network. In some use cases, the camera is aimed at the target rather than focusing on the shooter, and the camera and microphone pick up other shots without naturally determining that the shot was fired by the shooter of interest. Acoustic signatures are generated quickly and are often performed on a mobile device (e.g., a smartphone) that includes the camera and microphone. The mobile device is capable of executing instructions (e.g., applications) that include the components and systems described herein, so that classification and acoustic signature determination are performed on the mobile device. In some cases, the pre-trained network may be trained on a relatively comprehensive universe of shots. The learnable input layer may pre-transform the input data to achieve high-probability recognition by the pre-trained network of weapons and environments with specific munitions. This increases the likelihood of detecting shots of interest and rejecting all other shots.

[0055] According to some embodiments, the systems and methods described herein converge to useful weight masks where the gradient of the function computed by the DNN is monotonically increasing or monotonically decreasing at each input.

[0056] Referring to FIG. 2, FIG. 2 illustrates pre-training a DNN, where process 200 begins at 202, where an audio file is opened at block 204. The audio file may be the first audio file, or the next or subsequent audio file. Using the audio file, the system obtains blocks of samples framing the first and / or next shot at step 206. In other words, each shot is windowed with time-constrained samples. At block 208, spectrograms associated with the samples are generated and labeled with attributes.

[0057] In block 210, the system determines whether the most recent sample is associated with the last shot; if not, the system returns to block 206 to retrieve a block of samples associated with a subsequent shot. If so, the system proceeds to block 212 to determine whether the labeled spectrogram being created in block 208 is the last file. If not, the system returns to block 204 to open or retrieve the next audio file. If the system determines that the most recent file is the last file, the system proceeds to block 214 to aggregate the labeled spectrograms. In block 216, a DNN is trained on the labeled spectrograms. The system stops with the trained DNN in block 218.

[0058] Referring to FIG. 3, which illustrates deriving spectrogram weights, process 300 begins at 302, where a set of shots is acquired at block 304. The shots may be acquired by audio and / or video recording devices or may involve opening files associated with one or more shots. At block 306, the system acquires blocks of samples framing the first and / or next shots. In other words, each shot is windowed with time-bound samples. At block 308, spectrograms associated with the samples are created and labeled with attributes.

[0059] In block 310, the system determines whether the most recent spectrogram is associated with the last shot; if not, the system returns to block 306 to obtain a block of samples associated with a subsequent shot. If so, the system proceeds to block 312 to determine whether the labeled spectrogram created in block 308 is the last set. If not, the system returns to block 304 to obtain the next set of shots. If the system determines that the most recent file is the last file, the system proceeds to block 314, where the spectrograms to be labeled are aggregated. In block 316, the system updates W (weighting) until the value converges. The system stops with the spectrogram weighting in block 218.

[0060] Referring to FIG. 4, a process for classifying a shot 400 is described. The process begins at block 402, where, at block 404, the system obtains a block of samples framing the shot. At block 406, the system determines a spectrogram associated with the block of samples. At block 408, the system determines the Hadamard product of the spectrogram and a weight matrix. At block 410, the Hadamard product of the spectrogram and the weight matrix is ​​applied to the DNN. At block 412, the system makes a binary decision as to whether the shot is associated with the signature of the firearm in question or not. In some cases, the system can identify the type of firearm and the ammunition fired through the firearm. For example, when receiving an audio sample, the system can determine that the acoustic signature of the firearm in the audio sample matches a 230-grain round-nose bullet fired from a Beretta .45 ACP, without prior knowledge of the firearm in question. The process ends at block 414.

[0061] The system may include one or more processors and one or more computer-readable media that may store various modules, applications, programs, or other data. The computer-readable media may include instructions that, when executed by the one or more processors, cause the processors to perform the operations described herein for the system.

[0062] In some embodiments, a processor may include a central processing unit (CPU), a graphical processing unit (GPU), both a CPU and a GPU, a microprocessor, a digital signal processor, or other processing units or components well known in the art. Alternatively, or in addition, the functions described herein may be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), etc. Furthermore, each processor may have its own local memory. The local memory may store program modules, program data, and / or one or more operating systems. One or more control systems, computer controllers, and remote controls may include one or more cores.

[0063] Embodiments may be provided as a computer program product including a non-transitory machine-readable storage medium having stored thereon instructions (in compressed or uncompressed format) that can be used to program a computer (or other electronic device) to perform the steps or methods described herein. Computer-readable media may include volatile and / or non-volatile memory, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Machine-readable storage media may include, but are not limited to, hard drives, floppy disks, optical disks, CD-ROMs, DVDs, read-only memory (ROM), random-access memory (RAM), EPROMs, EEPROMs, flash memory, magnetic or optical cards, solid-state memory devices, or other types of media / machine-readable media suitable for storing electronic instructions. Furthermore, embodiments may be provided as a computer program product including a transitory machine-readable signal (in compressed or uncompressed format). Examples of machine-readable signals include, but are not limited to, signals that can be configured to be accessed by a computer system or machine that hosts or executes a computer program, including signals downloaded over the Internet or other networks, whether or not modulated using a carrier wave. In some examples, the process is performed on a mobile electronic device, such as a smartphone, tablet, laptop, or other mobile computing device. For example, a smartphone can download and execute an application (app) that utilizes a built-in microphone and / or video camera to capture sounds and / or audio-video associated with a firearm, perform the methods described herein to classify the shot, and determine whether the shot was associated with a particular firearm and ammunition combination.

[0064] Those skilled in the art will recognize that the processes or methods disclosed herein can be varied in many ways. The process parameters and order of the processes described and / or illustrated herein are provided by way of example only and can be varied as desired. For example, although the processes illustrated and / or described herein may be shown or discussed in a particular order, these processes do not necessarily have to be performed in the order shown or discussed.

[0065] The various exemplary methods described and / or illustrated herein may also omit one or more of the steps described or illustrated herein or may include additional steps in addition to those disclosed. Furthermore, the steps of any method disclosed herein may be combined with any one or more steps of any other method disclosed herein.

[0066] Since the present disclosure illustrates exemplary embodiments, it is not intended to limit the scope of the embodiments of the present disclosure and the appended claims in any way. The embodiments have been described above using functional blocks illustrating the implementation of specified functions and their relationships. The boundaries of these functional blocks have been arbitrarily defined herein for the convenience of description. Alternative boundaries may be defined as long as the specified functions and their relationships are appropriately performed.

[0067] The foregoing description of specific embodiments generally clarifies the general nature of the embodiments of the present disclosure, so that others, by applying the knowledge of those skilled in the art, can readily modify and / or adapt such specific embodiments for various uses without undue experimentation and without departing from the general concepts of the embodiments of the present disclosure. Therefore, such adaptations and modifications are intended to be within the scope and range of equivalents of the disclosed embodiments, based on the teaching and guidance presented herein. The phrases or terminology used herein are for the purpose of description and not of limitation, as would be interpreted by one of ordinary skill in the relevant art in light of the teaching and guidance presented herein.

[0068] The breadth and scope of the presently disclosed embodiments should not be limited by any of the above-described exemplary embodiments, but should be defined only in accordance with the following claims and their equivalents.

[0069] Conditional language such as "can," "could," "may," or "might," unless specifically stated otherwise or understood otherwise within the context in which it is used, is generally intended to convey that certain implementations may include certain features, elements, and / or operations, but not other implementations. Thus, such conditional language is generally not intended to imply that features, elements, and / or operations are required in one or more implementations, or necessarily include logic for one or more implementations to determine whether those features, elements, and / or operations are included in or performed in any particular implementation, with or without user input or prompting.

[0070] Unless otherwise noted, the terms "connected to" and "coupled to" (and their derivatives) as used herein shall be interpreted as allowing for both direct and indirect (i.e., via other elements or components) connections. Additionally, the terms "a" or "an" as used herein shall be interpreted as meaning "at least one." Finally, for ease of use, the terms "including" and "having" (and their derivatives), as used herein, shall not exclude additional components but shall be interpreted as open-ended.

[0071] This specification and the accompanying drawings disclose examples of systems, instruments, devices, and techniques that may provide systems and methods for determining the acoustic signature of a fired firearm. Of course, it is not possible to describe every conceivable combination of elements and / or methods for purposes of describing the various features of the present disclosure, but those skilled in the art will recognize that many further combinations and permutations of the disclosed features are possible. Accordingly, various modifications can be made to the present disclosure without departing from the scope or spirit of the present disclosure. Moreover, other embodiments of the present disclosure will become apparent from consideration of this specification and the accompanying drawings, and the practice of the disclosed embodiments set forth herein. The examples presented in this specification and the accompanying drawings are to be considered in all respects as illustrative and not restrictive. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.

[0072] Those skilled in the art will understand that in some implementations, the functionality provided by the processes and systems described above may be provided in alternative ways, such as being split into more software programs or routines or consolidated into fewer programs or routines. Similarly, in some implementations, the illustrated processes and systems may provide more or fewer functions than described, such as when other illustrated processes alternatively lack or include such functions, respectively, or when the amount of functionality provided is changed. Also, while various operations may be illustrated as being performed in a particular manner (e.g., serially or in parallel) and / or in a particular order, those skilled in the art will understand that in other implementations, the operations may be performed in other orders and manners. Those skilled in the art will understand that the data structures described above may be structured differently, such as by splitting a single data structure into multiple data structures or by aggregating multiple data structures into a single data structure. Similarly, in some implementations, the illustrated data structures may store more or less information than described, such as when other illustrated data structures alternatively lack or include such information, respectively, or when the amount or type of information stored is changed. The various methods and systems as illustrated in the figures and described herein represent example implementations. In other implementations, the methods and systems may be implemented in software, hardware, or a combination thereof. Similarly, in other implementations, the order of any method may be changed, and various elements may be added, reordered, combined, omitted, modified, etc.

[0073] From the foregoing, it will be understood that, while particular implementations have been described herein for purposes of illustration, various modifications can be made without departing from the spirit and scope of the appended claims and the elements described in the specification. Also, while particular embodiments are set forth below in particular claim forms, the inventors contemplate various embodiments in any available claim form. For example, while only some aspects may currently be described as embodied in a particular configuration, other aspects may likewise be so embodied. Various modifications and changes can be made, as will be apparent to those skilled in the art having the benefit of this disclosure. All such modifications and changes are intended to be included, and therefore, the foregoing description is to be regarded in an illustrative rather than a limiting sense. [Brief explanation of the drawings]

[0074] [Figure 1] A sample system for determining an acoustic signature according to some embodiments is described. [Figure 2] 1 illustrates a sample flowchart for pre-training a DNN, according to some embodiments. [Figure 3] 1 illustrates a sample flowchart for deriving spectrogram weightings, according to some embodiments. [Figure 4] 1 illustrates a sample flowchart for classifying and determining acoustic signatures.

Claims

1. 1. A method for determining an acoustic signature of a firearm, comprising: receiving an audio signal associated with the firing of a firearm; generating a spectrogram associated with the audio signal; applying a mask to the spectrogram to generate a masked spectrogram input; running a deep neural network configured to generate a classification of the masked spectrogram input; and determining a firearm and ammunition combination associated with the audio signal based at least in part on the masked spectrogram input.

2. The method of claim 1 , wherein the step of receiving an audio signal is performed by a microphone associated with a smartphone.

3. The method of claim 2 , wherein the method is performed on the smartphone.

4. The method of claim 1 , further comprising training the mask with training data associated with the firearm and the ammunition.

5. The method of claim 4 , wherein the mask is a Hadamard product of a weight matrix and an input vector.

6. receiving subsequent audio data associated with the firing of the second firearm; 10. The method of claim 1, further comprising determining that the shot was fired from a second firearm not associated with the firearm and ammunition combination associated with the audio signal.

7. 10. The method of claim 1, further comprising receiving a plurality of audio samples associated with the discharge of a plurality of firearms, and identifying, from the plurality of audio samples, individual ones of the plurality of audio samples associated with the combination of the firearm and the ammunition.

8. 1. A system for determining an acoustic signature of a firearm, comprising a computing device having one or more processors, the one or more processors configured with instructions, comprising: The computing device a spectrogram module configured to receive an audio signal and output a spectrogram associated with the audio signal; an attribute labeler configured to label the spectrogram associated with the audio signal with attributes to generate a labeled spectrogram; a weighting module configured to aggregate and weight the labeled spectrograms to generate a weighted spectrogram; a deep neural network classifier configured to classify the weighted spectrogram to determine a firearm associated with the audio signal.

9. The system of claim 8 , wherein the computing device is a smartphone.

10. The system of claim 8 , further comprising a microphone configured to capture the audio signal.

11. The system of claim 8 , further comprising a camera configured to capture audio / video signals.

12. 9. The system of claim 8, wherein the weighting module is further configured to learn a weight mask on the spectrogram to generate an acoustic signature of the firearm.

13. The system of claim 12 , wherein the weight mask is generated by determining a Hadamard product of a weight matrix and an input vector to generate the weighted spectrogram.

14. 9. The system of claim 8, wherein the system is configured to distinguish between a first acoustic signature associated with a first firearm and a second acoustic signature associated with a second firearm.