A system and method for threat detection using real time audio recording
Patent Information
- Application Number
- IN202341079079
- Authority / Receiving Office
- IN · IN
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2023-11-21
- Publication Date
- 2026-08-07
- Estimated Expiration
- 2043-11-21
AI Technical Summary
Existing live audio event detection systems require external hardware modules for audio signal detection, making them user-unfriendly and challenging to maintain, and are limited to pre-recorded audio data, lacking real-time threat prediction capabilities.
A system and method for real-time audio threat detection using deep learning-based multi-class classification, incorporating signal analysis techniques like Fast Fourier transform and decibel spectrogram, which identifies potential threats without external hardware and sends alerts through communication platforms like email and SMS, utilizing preprocessor modules for data cleaning and prediction with customizable deep learning models.
Enables user-friendly, hardware-free real-time audio threat detection and prediction, sending immediate alerts based on signal intensity and frequency analysis, improving the accuracy and range of threat detection by using customizable deep learning models and integrating audio recording with prediction modules.
Abstract
Description
TECHNICAL FIELD
[0001] The present invention relates to the field of audio event detection. Inparticular, it relates to a system and method for threat detection using real timeaudio recording, together with the concept of audio analysis and classification.BACKGROUND
[0002] Event detection consist of the feature extraction of each signal and theclassification model considering the feature vector as input. Event detection canbe achieved through various types of technique based on audio signal, image,video and more. The audio signal type is a robust signal in comparison to thevisual changing, and also, as they are collected non-contact, they are highly usefulto the field of event detection. In recent years, audio event detection (AED) vastlydeveloped as an important field due to several emergent applications in the fieldof smart home automation, human activity detection, elderly and child caresystem and sound anomaly detection.
[0003] One of the existing applications US20180268674A1 titled "Dynamicidentification of threat level associated with a person using an audio / videorecording and communication device" discloses a dynamic identification of threatlevel associated with a person using an audio / video recording and communicationdevice. The existing application provides prediction based on video and which istime consuming, and accuracy of prediction varies with respect to cameraposition. To overcome the disadvantage of said feature, usually, machine learningor deep learning models are used for both of a feature extraction and aclassification in audio-based event detection. A. A. Al-Tameem et. al., in theirwork entitled "machine learning approach for identification of threat content inaudio messages shared on social media" discussed about the detection of unusualcontent in the audio messages shared in social media for identifying hazards orplanned threats. The main problem with the existing models is they detect andclassify pre-recorded audio data alone and identify only the audio data.
[0004] However, live event audio detection is highly needed in the currentscenario which demands live audio recording modules and real time threatprediction. The problem with existing live audio event detection systemsavailable in the market is that they require external hardware modules for audiodetection, which include smart wearable devices, or unmanned aerial vehicles etc.The existing products are not user-friendly as the maintenance or usage of theseparate hardware module might pose a challenge to regular usage.
[0005] Hence, there is a need for system and method to overcome the use ofexternal hardware module for audio signal detection in a live audio eventdetection platform.OBJECTS OF THE PRESENT DISCLOSURE
[0006] Some of the objects of the present disclosure, which at least oneembodiment herein satisfies are as listed herein below.
[0007] It is an object of the present disclosure to provide a system and method forthreat detection using real time audio recording.
[0008] It is another object of the present invention to provide a system andmethod threat detection using real time audio recording, which identify the type ofreal-world audio data using deep learning based multi-class classificationtechnique(s), and determine whether the identified sound type corresponds to apotentially threatful environment.
[0009] It is another object of the present invention to provide a system andmethod for threat detection using real time audio recording, which implementsound classification without incorporating a hardware module.
[0010] It is another object of the present invention to provide a system andmethod for threat detection using real time audio recording, which uses signalanalysis technique such as Fast Fourier transform and decibel spectrogram toanalyze the signal changes over time from the recorded audio which is the inputfor the live detection system and based on the signal intensity and frequency ofthe sound wave, severity of the incident prediction is carried out.
[0011] It is another object of the present invention to provide a system andmethod for threat detection using real time audio recording, which send animmediate alert message automatically to the registered help servicing departmentthrough their emergency contact via e-mail, Short Message Service (SMS), andWhatsApp which also contains the real-time audio recording file as an attachment.
[0012] It is another object of the present invention to provide a system andmethod for threat detection using real time audio recording, which uses anintegration of the audio recording module with the prediction module, and thefunctionality of automatically prompting the prediction module to identify thetype of sound detected, and sending the recorded sound to the registered help as athreat alert system, asking for their immediate assistance.
[0013] It is another object of the present invention to provide a system andmethod for threat detection using real time audio recording, which adopts threedifferent deep-learning modules and the user can choose the appropriate modelbased on the accuracy achieved during training for threat detection, which mightdiffer from model to model based on the training dataset used.
[0014] It is another object of the present invention to provide a system andmethod for threat detection using real time audio recording, in which the dataset iscustomizable so that many more classes of audio data can be added to train themodels, so as to improve the range and robustness of the application.SUMMARY
[0015] The present invention relates to the field of audio event detection. Inparticular, it relates to a system and method for threat detection using real timeaudio recording.
[0016] An aspect of the present disclosure pertains to a system for threat detectionusing real time audio recording. The system comprising: one or more preprocessor, a memory coupled to the one or more processor. The memorycomprises processor-executable instructions to cause the one or more preprocessors to: receive at least one data packet from a computing device associatedwith one or more user, the at least one data packet pertains to one or more audiodata captured from an audio recording module. Further, the system can beconfigured to perform a data-cleaning process for generating a mono-channelaudio data based on the one or more audio data, and execute a training process onthe one or more audio data to obtain a trained mono-channel audio data. Further,the system can be configured to record the one or more audio data in real time byan audio detection module to identify one or more incidents associated with theone or more user, and store the one or more audio data in a pre-defined format.Furthermore, the system can be configured to identify one or more classassociated with the one or more audio data by a prediction module, and detect theone or more audio data as a suspicious event by one or more deep learning modelselected by the one or more user based on the trained mono-channel audio data.Finally, the system can be configured to generate and transmit an alert message toone or more registered contact list through one or more communication platformsbased on detection of suspicious event associated with the one or more audio data.
[0017] In an aspect, the audio recording module is configured to: execute one ormore operations pertaining to the one or more audio data in real time, wherein theone or more operations comprises at least one of a detect audio, a record audio, astop recording, and a save recording.
[0018] In an aspect, the one or more audio data comprises at least one of a speechdata, an object sound, a music recording, and a digitally recorded human sounds;the one or more incidents comprises at least one of a gunshot, a scream, a hittingsound, a fire alarm and an explosion.
[0019] In an aspect, the data cleaning process comprises at least one of a samplingdown the one or more audio data to at least 16000 Hz mono-channel audio datafrom an original frequency.
[0020] In an aspect, the training process comprises train the mono-channel audiodata using the one or more deep-learning model and store the trained the monochannel audio data.
[0021] In an aspect, the one or more deep learning model comprises at least oneof a long-short term memory (LSTP), a 1-dimensional convolutional neuralnetwork (1D CNN), and a 2-dimensional convolutional neural network (2DCNN).
[0022] In an aspect, the pre-defined format comprises at least one of a wave formfile format, an audio interchange file format, an audio file format, and a pulsecode modulation file format.
[0023] In an aspect, the registered contact list comprises at least one of a policestation, a pink police patrol, and a cyber cell.
[0024] In an aspect, the communication platform comprises at least one of an email, a Short Message Service (SMS) and a WhatsApp.
[0025] In an aspect, a method for threat detection using real time audio recording.The method includes steps of receiving, by a system, at least one data packet fromcomputing device associated with one or more users. The at least one data packetpertains to one or more audio data captured from an audio recording module. Themethod includes steps of performing, by the system, a data-cleaning process forgenerating a mono-channel audio data based on the one or more audio data, andexecute a training process on the one or more audio data to obtain a trained monochannel audio data. Further, the method comprises the step of recording, by thesystem, the one or more audio data in real time by an audio detection module toidentify one or more incidents associated with the one or more user, and store theone or more audio data in a pre-defined format. Further, the method comprises thestep of identifying, by the system (102), one or more class associated with the oneor more audio data by a prediction module, and detect the one or more audio dataas a suspicious event by one or more deep learning model selected by the one ormore user based on the trained mono-channel audio data. Finally, generating andtransmitting, by the system, an alert message to one or more registered contact listthrough one or more communication platforms based on detection of suspiciousevent associated with the one or more audio data.
[0026] Various objects, features, aspects, and advantages of the inventive subjectmatter will become more apparent from the following detailed description ofpreferred embodiments, along with the accompanying drawing figures in whichlike numerals represent like components.BRIEF DESCRIPTION OF DRAWINGS
[0027] The accompanying drawings are included to provide a furtherunderstanding of the present disclosure, and are incorporated in, and constitute apart of this specification. The drawings illustrate exemplary embodiments of thepresent disclosure, and together with the description, serve to explain theprinciples of the present disclosure.
[0028] In the figures, similar components, and / or features may have the samereference label. Further, various components of the same type may bedistinguished by following the reference label with a second label thatdistinguishes among the similar components. If only the first reference label isused in the specification, the description is applicable to any one of the similarcomponents having the same first reference label irrespective of the secondreference label.
[0029] FIG. 1 illustrates exemplary network architecture of the proposed systemfor threat detection using real time audio recording, in accordance with anembodiment of the present disclosure.
[0030] FIG. 2 illustrates architecture of the proposed system for threat detectionusing real time audio recording, in accordance with an embodiment of the presentdisclosure.
[0031] FIG. 3 illustrates an exemplary view of a flow diagram of the proposedmethod for threat detection using real time audio recording, in accordance with anembodiment of the present disclosure.
[0032] FIG.4 illustrates an exemplary computer system in which or with whichembodiments of the present invention can be utilized in accordance withembodiments of the present disclosure.DETAILED DESCRIPTION
[0033] The following is a detailed description of embodiments of the disclosuredepicted in the accompanying drawings. The embodiments are in such detail as toclearly communicate the disclosure. However, the amount of detail offered is notintended to limit the anticipated variations of embodiments; on the contrary, theintention is to cover all modifications, equivalents, and alternatives falling withinthe spirit and scope of the present disclosure as defined by the appended claims.
[0034] Various aspects of the present disclosure are described with respect toFIGs. 1-4.
[0035] The present invention relates to the field of audio event detection. Inparticular, it relates to a system and method for threat detection using real timeaudio recording.
[0036] An aspect of the present disclosure pertains to a system for threat detectionusing real time audio recording. The system comprising: one or more preprocessor, a memory coupled to the one or more processor. The memorycomprises processor-executable instructions to cause the one or more preprocessors to: receive at least one data packet from a computing device associatedwith one or more user, the at least one data packet pertains to one or more audiodata captured from an audio recording module. Further, the system can beconfigured to perform a data-cleaning process for generating a mono-channelaudio data based on the one or more audio data, and execute a training process onthe one or more audio data to obtain a trained mono-channel audio data. Further,the system can be configured to record the one or more audio data in real time byan audio detection module to identify one or more incidents associated with theone or more user, and store the one or more audio data in a pre-defined format.Furthermore, the system can be configured to identify one or more classassociated with the one or more audio data by a prediction module, and detect theone or more audio data as a suspicious event by one or more deep learning modelselected by the one or more user based on the trained mono-channel audio data.Finally, the system can be configured to generate and transmit an alert message toone or more registered contact list through one or more communication platformsbased on detection of suspicious event associated with the one or more audio data.
[0037] In an aspect, the audio recording module is configured to: execute one ormore operations pertaining to the one or more audio data in real time, wherein theone or more operations comprises at least one of a detect audio, a record audio, astop recording, and a save recording.
[0038] In an aspect, the one or more audio data comprises at least one of a speechdata, an object sound, a music recording, and a digitally recorded human sounds;the one or more incidents comprises at least one of a gunshot, a scream, a hittingsound, a fire alarm and an explosion.
[0039] In an aspect, the data cleaning process comprises at least one of a samplingdown the one or more audio data to at least 16000 Hz mono-channel audio datafrom an original frequency.
[0040] In an aspect, the training process comprises train the mono-channel audiodata using the one or more deep-learning model and store the trained the monochannel audio data.
[0041] In an aspect, the one or more deep learning model comprises at least oneof a long-short term memory (LSTP), a 1-dimensional convolutional neuralnetwork (1D CNN), and a 2-dimensional convolutional neural network (2DCNN).
[0042] In an aspect, the pre-defined format comprises at least one of a wave formfile format, an audio interchange file format, an audio file format, and a pulsecode modulation file format.
[0043] In an aspect, the registered contact list comprises at least one of a policestation, a pink police patrol, and a cyber cell.
[0044] In an aspect, the communication platform comprises at least one of an email, a Short Message Service (SMS) and a WhatsApp.
[0045] In an aspect, a method for threat detection using real time audio recording.The method includes steps of receiving, by a system, at least one data packet fromcomputing device associated with one or more users. The at least one data packetpertains to one or more audio data captured from an audio recording module. Themethod includes steps of performing, by the system, a data-cleaning process forgenerating a mono-channel audio data based on the one or more audio data, andexecute a training process on the one or more audio data to obtain a trained monochannel audio data. Further, the method comprises the step of recording, by thesystem, the one or more audio data in real time by an audio detection module toidentify one or more incidents associated with the one or more user, and store theone or more audio data in a pre-defined format. Further, the method comprises thestep of identifying, by the system (102), one or more class associated with the oneor more audio data by a prediction module, and detect the one or more audio dataas a suspicious event by one or more deep learning model selected by the one ormore user based on the trained mono-channel audio data. Finally, generating andtransmitting, by the system, an alert message to one or more registered contact listthrough one or more communication platforms based on detection of suspiciousevent associated with the one or more audio data.
[0046] FIG. 1 illustrates exemplary network architecture 100 of the proposedsystem 102 for threat detection using real time audio recording, in accordancewith an embodiment of the present disclosure.
[0047] In an embodiment, referring to FIG. 1, the system 102 will be connected toa network 104, which is further connected to at least one computing devices 108-1, 108-2, … 108-N (collectively referred as computing device 108, herein)associated with one or more users devices 106-1, 106-2, … 106-N (collectivelyreferred as user 106, herein). The computing device 108 may be personalcomputers, laptops, tablets, wristwatch or any custom-built computing deviceintegrated within a modern diagnostic machine that can connect to a network asan IoT (Internet of Things) device. Furthermore, the network 104 can beconfigured with a centralized server 110 that stores compiled data from all thesecure transactions.
[0048] In an embodiment, the system 102 may receive at least one input data fromthe at least one computing devices 108. A person of ordinary skill in the art willunderstand that the at least one computing devices 108 may be individuallyreferred to as computing device 108 and collectively referred to as computingdevices 108. In an embodiment, the computing device 110 may also be referred toas User Equipment (UE). Accordingly, the terms "computing device" and "UserEquipment" may be used interchangeably throughout the disclosure.
[0049] In an embodiment, the computing device 108 may transmit the at least onereceived data packet over a point-to-point or point-to-multipoint communicationchannel or network 104 to the system 102.
[0050] In an embodiment, the computing device 108 may involve collection,analysis, and sharing of data received from the system 102 via the communicationnetwork 104.
[0051] In an embodiment, the system 102 may execute one or more instructionfor threat detection using real time audio recording.
[0052] In an exemplary embodiment, the system 102 may include, but not belimited to, a computer enabled device, a mobile phone, a tablet, a display device, adisplay projector, a AR / VR / MR, a imaging device, a sensors, a NFC, a network(Wired or Wireless), an apparatus to dispatch gift, prints, ecommerce,instructions, a Remote Detection Service (Detection Device) enabled devices suchas iBeacon technologies, NFC, IR / RF services, bluetooth to detect the devicesnearby, a connect signs objects, an apparatus, a vending machine, a gift clawmachine, a combination of the vending machine, and the gift claw machine, adrone, a robot, an advertisement displays, or some combination thereof.
[0053] In an exemplary embodiment, the communication network 104 mayinclude, but not be limited to, at least a portion of one or more networks havingone or more nodes that transmit, receive, forward, generate, buffer, store, route,switch, process, or a combination thereof, etc. one or more messages, packets,signals, waves, voltage or current levels, some combination thereof, or so forth. Inan exemplary embodiment, the communication network 104 may include, but notbe limited to, a wireless network, a wired network, an internet, an intranet, apublic network, a private network, a packet-switched network, a circuit-switchednetwork, an ad hoc network, an infrastructure network, a Public-SwitchedTelephone Network (PSTN), a cable network, a cellular network, a satellitenetwork, a fiber optic network, or some combination thereof.
[0054] In an embodiment, the one or more computing devices 110 maycommunicate with the system 102 via a set of executable instructions residing onany operating system. In an embodiment, the one or more computing devices 110may include, but not be limited to, any electrical, electronic, electro-mechanical,or an equipment, or a combination of one or more of the above devices such asmobile phone, smartphone, Virtual Reality (VR) devices, Augmented Reality(AR) devices, laptop, a general-purpose computer, desktop, personal digitalassistant, tablet computer, mainframe computer, or any other computing device,wherein the one or more computing devices 110 may include one or more in-builtor externally coupled accessories including, but not limited to, a visual aid devicesuch as imaging device, audio aid, a microphone, a keyboard, input devices suchas touch pad, touch enabled screen, electronic pen, receiving devices for receivingany audio or visual signal in any range of frequencies, and transmitting devicesthat can transmit any audio or visual signal in any range of frequencies. It may beappreciated that the one or more computing devices 110 may not be restricted tothe mentioned devices and various other devices may be used.
[0055] In an embodiment, the network 104 is further configured with acentralized server 110 including a database, where the user identity data is usedfor providing authentication to the users. It can be retrieved based on therequirement.
[0056] In an embodiment, the system 102 can be configured to receive at leastone data packet from computing device associated with one or more users. Thus, afirst connection can be established between the system 102 and the user 106associated with the computing device 108. The at least one data packet pertains toone or more audio data captured from an audio recording module and the audiodata comprises at least one of a speech data, an object sound, a music recording,and a digitally recorded human sounds.
[0057] In an embodiment, the system 102 can be configured to perform a datacleaning process for generating a mono-channel audio data based on the one ormore audio data, and execute a training process on the one or more audio data toobtain a trained mono-channel audio data.
[0058] In an embodiment, the system 102 can be configured to record the one ormore audio data in real time by an audio detection module to identify one or moreincidents associated with the one or more user, and store the one or more audiodata in a pre-defined format.
[0059] In an embodiment, the system 102 can be configured to identify one ormore class associated with the one or more audio data by a prediction module, anddetect the one or more audio data as a suspicious event by one or more deeplearning model selected by the one or more user based on the trained monochannel audio data.
[0060] In an embodiment, the system 102 can be configured to generate andtransmit an alert message to one or more registered contact list through one ormore communication platforms based on detection of suspicious event associatedwith the one or more audio data.
[0061] Although FIG. 1 shows exemplary components of the network architecture100, in other embodiments, the network architecture 100 may include fewercomponents, different components, differently arranged components, or additionalfunctional components than depicted in FIG. 1. Additionally, or alternatively, oneor more components of the network architecture 100 may perform functionsdescribed as being performed by one or more other components of the networkarchitecture 100.
[0062] FIG. 2 illustrates architecture (200) of the proposed system for threatdetection using real time audio recording.
[0063] In an aspect, referring to FIG. 2, the system 102 may comprise one ormore processor(s) 202. The one or more processor(s) 202 may be implemented asone or more microprocessors, microcomputers, microcontrollers, edge or fogmicrocontrollers, digital signal processors, central processing units, logiccircuitries, and / or any devices that process data based on operational instructions.Among other capabilities, the one or more processor(s) 202 may be configured tofetch and execute computer-readable instructions stored in a memory 204 of thesystem 102. The memory 204 may be configured to store one or more computerreadable instructions or routines in a non-transitory computer readable storagemedium, which may be fetched and executed to create or share data packets over anetwork service. The memory 204 may comprise any non-transitory storagedevice including, for example, volatile memory such as Random Access Memory(RAM), or non-volatile memory such as Erasable Programmable Read-OnlyMemory (EPROM), flash memory, and the like.
[0064] Referring to FIG. 2, the system 102 may include an interface(s) 206. Theinterface(s) 206 may comprise a variety of interfaces, for example, interfaces fordata input and output devices, referred to as I / O devices, storage devices, and thelike. The interface(s) 206 may facilitate communication to / from the system 102.The interface(s) 206 may also provide a communication pathway for one or morecomponents of the system 102. Examples of such components include, but are notlimited to, processing unit / engine(s) 208 and a local database 210.
[0065] In an embodiment, the processing unit / engine(s) 208 may be implementedas a combination of hardware and programming (for example, programmableinstructions) to implement one or more functionalities of the processing engine(s)208. In examples described herein, such combinations of hardware andprogramming may be implemented in several different ways. For example, theprogramming for the processing engine(s) 208 may be processor-executableinstructions stored on a non-transitory machine-readable storage medium and thehardware for the processing engine(s) 208 may comprise a processing resource(for example, one or more processors), to execute such instructions. In the presentexamples, the machine-readable storage medium may store instructions that, whenexecuted by the processing resource, implement the processing engine(s) 208. Insuch examples, the system 102 may comprise the machine-readable storagemedium storing the instructions and the processing resource to execute theinstructions, or the machine-readable storage medium may be separate butaccessible to the system 102 and the processing resource. In other examples, theprocessing engine(s) 208 may be implemented by electronic circuitry.
[0066] In an embodiment, the local database 210 may comprise data that may beeither stored or generated as a result of functionalities implemented by any of thecomponents of the processor 202 or the processing engines 208. In anembodiment, the local database 210 may be separate from the system 102.
[0067] In an exemplary embodiment, the processing engine 208 may include oneor more engines selected from any of an audio data acquisition module 212, anaudio data cleaning module 214, a live audio data recording module 216, an audiodata identification and detection module 218, an alert generation module 220, andother modules 222 having functions that may include but are not limited totesting, storage, and peripheral functions, such as wireless communication unit forremote operation, audio unit for alerts and the like.
[0068] In an embodiment, the data acquisition module 212 may include meansreceiving at least one data packet from the users 106 associated with thecomputing devices 108, the at least one data packet pertains to one or more audiodata retrieved from the users 106 which comprises at least one of a speech data, anobject sound, a music recording, a digitally recorded human sounds, a far-flungspeech, and a background noise.
[0069] In an embodiment, the audio data cleaning module 214 may be configuredto process the one or more audio data to sample down the one or more audio datato a mono-channel audio data by data-cleaning. Further, train the mono-channelaudio data using one or more deep-learning model and store the trained the monochannel audio data for use in additional steps.
[0070] In an embodiment, the live audio data recording module 216 may beconfigured to record the one or more audio data in real time to identify one ormore incidents associated with the one or more user, and store the one or moreaudio data in a pre-defined format retrieved from the users 106. The audiorecording module is configured to execute one or more operations pertaining tothe one or more audio data in real time, wherein the one or more operationscomprises at least one of a detect audio, a record audio, a stop recording whenthere is no more the one or more audio data, and a save the one or more audio datarecording in the pre-defined format, and a prompt the prediction module toidentify the one or more class of the one or more audio data detected.
[0071] In an embodiment, the audio data identification and detection module 218is configured to identify one or more class associated with the one or more audiodata by a prediction module, and detect the one or more audio data as a suspiciousevent by one or more deep learning model selected by the one or more user basedon the trained mono-channel audio data.
[0072] In an embodiment, the alert generating module 218 may be configured togenerate and transmit an alert message to one or more registered contact listthrough one or more communication platforms based on detection of suspiciousevent associated with the one or more audio data.
[0073] In an embodiment, signal analysis technique such as Fast Fouriertransform and decibel spectrogram is used to analyse the signal changes over timefrom the recorded audio data which is the input for the live detection system andbased on the signal intensity and frequency of the audio data, severity of theincident prediction is carried out.
[0074] FIG. 3 illustrates view of a flow diagram of the proposed method for threatdetection using real time audio recording, in accordance with an embodiment ofthe present disclosure.
[0075] In an embodiment, referring to FIG.3, the proposed method 300 forfacilitating detection using real time audio recording. At step 302 the systemstores a dataset with one or more audio data belongs to one or more classes. Theone or more audio data comprises at least one of a speech data, an object sound, amusic recording, and a digitally recorded human sounds, a far-flung speech, and abackground noise. Further, at step 304 the one or more audio data received fromuser 106 is processed through a data-cleaning process. The data cleaning processcomprises at least one of a sampling down the one or more audio data to at least16000 Hz mono-channel audio data from an original frequency. The one or moredeep learning model comprises at least one of a long-short term memory (LSTP),a 1-dimensional convolutional neural network (1D CNN), and a 2-dimensionalconvolutional neural network (2D CNN). At step 308, the trained mono-channelaudio data is stored separately for using during the prediction step. At step 310 theone or more user can select one or more deep learning model for using in theprediction step based on the training process on one or more audio data.
[0076] In an embodiment, referring to method 300, at steps 312 an audiodetection module records one or more audio in real-time to identify one or moreincidents associated with the one or more user. The audio detecting modulecomprises at least one of a microphone, a digital voice recorder, and the likes. Theone or more audio data comprises at least one of a speech data, an object sound, amusic recording, a digitally recorded human sounds, a far-flung speech, and abackground noise. The one or more incidents comprises at least one of a gun shot,a scream, a hitting sound, a fire alarm and an explosion. Further, at step 314, adata-cleaning process is performed for generating a mono-channel audio databased on the one or more audio data. The one or more audio data in real time isprocessed by one or more operations comprising at least one of a normalization, atrimming silence in the audio on both ends and an adding at least an extra 0.5second of blank audio at both ends. Further, at step 316, the pre-processed one ormore audio data is stored in a pre-defined format. The pre-defined formatcomprises at least one of a wave form file format, an audio interchange fileformat, an audio file format, and a pulse code modulation file format.
[0077] In an embodiment, referring to method 300, at step 318, a predictionmodule identifies one or more class associated with the one or more audio data,which is recorded in real time and stored in a pre-defined format, using one ormore deep learning model selected by the one or more user. The predictionmodule is configured to identify the one or more class of the one or more audiodata in real time using the one or more deep learning model, and determinewhether the identified the one or more audio data corresponds to a potentiallydangerous environment. The one or more deep learning model comprises at leastone of a long-short term memory (LSTP), a 1-dimensional convolutional neuralnetwork (1D CNN), and a 2-dimensional convolutional neural network (2DCNN). Further, at step 320, detecting the one or more audio data as a suspiciousevent by one or more deep learning model selected by the one or more user basedon the trained mono-channel audio data. Further, at step 322, generating andtransmitting an alert message to one or more registered contact list through one ormore communication platforms based on detection of suspicious event associatedwith the one or more audio data. The registered contact list comprises at least oneof a police station, a pink police patrol, a cyber cell, friends, family members andlocal guardian. The communication platform comprises at least one of an e-mail, aShort Message Service (SMS) and a WhatsApp. Further, the e-mail message alsocontains the audio file as an attachment, which was generated as a result of thereal-time audio recording. Further, at step 320, if the one or more audio data is notdetected as a suspicious event by one or more deep learning model selected by theone or more user, based on the trained mono-channel audio data, a predictionresult is shown at the output in step 324 and the process stops.
[0078] FIG.4 illustrates an exemplary computer system in which or with whichembodiments of the present invention can be utilized in accordance withembodiments of the present disclosure.
[0079] Referring to FIG. 4, computer system includes an external storage device410, a bus 420, a main memory 430, a read only memory 440, a mass storagedevice 450, communication port 460, and a processor 470. A person skilled in theart will appreciate that computer system may include more than one processor andcommunication ports. Examples of processor 470 include, but are not limited to,an Intel Itanium or Itanium 2 processor(s), or AMD Opteron or AthlonMP processor(s), Motorola lines of processors, FortiSOC system on a chipprocessors or other future processors. Processor 470 may include various modulesassociated with embodiments of the present invention. Communication port 460can be any of an RS-232 port for use with a modem based dialup connection, a10 / 100 Ethernet port, a Gigabit or 10 Gigabit port using copper or fiber, a serialport, a parallel port, or other existing or future ports. Communication port 460may be chosen depending on a network, such a Local Area Network (LAN), WideArea Network (WAN), or any network to which computer system connects.
[0080] In an embodiment, the memory 430 can be Random Access Memory(RAM), or any other dynamic storage device commonly known in the art. Readonly memory 440 can be any static storage device(s) e.g., but not limited to, aProgrammable Read Only Memory (PROM) chips for storing static informatione.g., start-up or BIOS instructions for processor 470. Mass storage 460 may beany current or future mass storage solution, which can be used to storeinformation and / or instructions. Exemplary mass storage solutions include, but arenot limited to, Parallel Advanced Technology Attachment (PATA) or SerialAdvanced Technology Attachment (SATA) hard disk drives or solid-state drives(internal or external, e.g., having Universal Serial Bus (USB) and / or Firewireinterfaces), e.g. those available from Seagate (e.g., the Seagate Barracuda 7102family) or Hitachi (e.g., the Hitachi Deskstar 7K1000), one or more optical discs,Redundant Array of Independent Disks (RAID) storage, e.g. an array of disks(e.g., SATA arrays), available from various vendors including Dot Hill SystemsCorp., LaCie, Nexsan Technologies, Inc. and Enhance Technology, Inc.
[0081] In an embodiment, the bus 420 communicatively couples processor(s) 470with the other memory, storage and communication blocks. Bus 420 can be, e.g. aPeripheral Component Interconnect (PCI) / PCI Extended (PCI-X) bus, SmallComputer System Interface (SCSI), USB or the like, for connecting expansioncards, drives and other subsystems as well as other buses, such a front side bus(FSB), which connects processor 470 to software system.
[0082] In another embodiment, operator and administrative interfaces, e.g. adisplay, keyboard, and a cursor control device, may also be coupled to bus 420 tosupport direct operator interaction with computer system. Other operator andadministrative interfaces can be provided through network connections connectedthrough communication port 460. External storage device 410 can be any kind ofexternal hard-drives, floppy drives, IOMEGA Zip Drives, Compact Disc - ReadOnly Memory (CD-ROM), Compact Disc - Re-Writable (CD-RW), Digital VideoDisk - Read Only Memory (DVD-ROM). Components described above are meantonly to exemplify various possibilities. In no way should the aforementionedexemplary computer system limit the scope of the present disclosure.
[0083] If the specification states a component or feature "may", "can", "could", or"might" be included or have a characteristic, that particular component or featureis not required to be included or have the characteristic.
[0084] As used in the description herein and throughout the claims that follow,the meaning of "a," "an," and "the" includes plural reference unless the contextclearly dictates otherwise. Also, as used in the description herein, the meaning of"in" includes "in" and "on" unless the context clearly dictates otherwise.
[0085] It is to be appreciated by a person skilled in the art that while variousembodiments of the present disclosure have been elaborated for system andmethod for threat detection using real time audio recording. However, theteachings of the present disclosure are also applicable for other types ofapplications as well, and all such embodiments are well within the scope of thepresent disclosure. However, the system and method threat detection using realtime audio recording is also equally implementable in other industries as well, andall such embodiments are well within the scope of the present disclosure withoutany limitation.
[0086] Accordingly, the present disclosure provides a system and method forthreat detection using real time audio recording.
[0087] Moreover, in interpreting the specification, all terms should be interpretedin the broadest possible manner consistent with the context. In particular, theterms "comprises" and "comprising" should be interpreted as referring toelements, components, or steps in a non-exclusive manner, indicating that thereferenced elements, components, or steps may be present, or utilized, orcombined with other elements, components, or steps that are not expresslyreferenced. Where the specification claims refer to at least one of somethingselected from the group consisting of A, B, C….and N, the text should beinterpreted as requiring only one element from the group, not A plus N, or B plusN, etc.
[0088] While the foregoing describes various embodiments of the disclosure,other and further embodiments of the disclosure may be devised without departingfrom the basic scope thereof. The scope of the disclosure is determined by theclaims that follow. The disclosure is not limited to the described embodiments,versions or examples, which are included to enable a person having ordinary skillin the art to make and use the disclosure when combined with information andknowledge available to the person having ordinary skill in the art.ADVANTAGES OF THE PRESENT DISCLOSURE
[0089] The present disclosure provides a system and method, which helps in thedetection of threats with the use of live audio classification on audio recorded inreal time.
[0090] The present disclosure provides a system and method, which identify thetype of real-world audio data using deep learning based multi-class classificationtechnique(s), and determine whether the identified sound type corresponds to apotentially threatful environment.
[0091] The present disclosure provides a system and method, which implementsound classification without incorporating a hardware module.
[0092] The present disclosure provides a system and method, which incorporatesdifferent deep learning classification and prediction techniques and an audiorecording module for live audio event detection.
[0093] The present disclosure provides a system and method, which uses signalanalysis technique such as Fast Fourier transform and decibel spectrogram toanalyze the signal changes over time from the recorded audio which is the inputfor the live detection system and based on the signal intensity and frequency ofthe sound wave, severity of the incident prediction is carried out.
[0094] The present disclosure provides a system and method, which send animmediate alert message automatically to the registered help servicing departmentthrough their emergency contact via e-mail, Short Message Service (SMS), andWhatsApp which also contains the real-time audio recording file as an attachment.
[0095] The present disclosure provides a system and method, which uses anintegration of the audio recording module with the prediction module, and thefunctionality of automatically prompting the prediction module to identify thetype of sound detected, and sending the recorded sound to the registered help as athreat alert system, asking for their immediate assistance.
[0096] The present disclosure provides a system and method, which adopts threedifferent deep-learning modules and the user can choose the appropriate modelbased on the accuracy achieved during training for threat detection, which mightdiffer from model to model based on the training dataset used.
[0097] The present disclosure provides a system and method, in which the datasetis customizable so that many more classes (types) of audio data can be added totrain the models, so as to improve the range and robustness of the application.
Claims
1. A system (102) for threat detection using real time audio recording, the system (102) comprising: one or more processors (202); and a memory coupled to the one or more processors (202), wherein said memory (204) stores instructions which when executed by the one or more processors (202) cause the system (102) to: receive at least one data packet from one or more computing 10 device associated with one or more users (110), wherein the at least one data packet pertains to one or more audio data captured from an audio recording module; perform a data-cleaning process for generating a mono-channel audio data based on the one or more audio data, and execute a training process on the one or more audio data to obtain a trained mono-channel audio data; record the one or more audio data in real time by an audio detection module to identify one or more incidents associated with the one or more user, and store the one or more audio data in a pre-defined format; 20 identify one or more class associated with the one or more audio data by a prediction module, and detect the one or more audio data as a suspicious event by one or more deep learning model selected by the one or more user based on the trained mono-channel audio data; and generate and transmit an alert message to one or more registered contact list through one or more communication platforms based on detection of suspicious event associated with the one or more audio data.
2. The system (102) as claimed in claim 1, wherein the audio recording module is configured to: execute one or more operations pertaining to the one or more audio data in 30 real time, wherein the one or more operations comprises at least one of a detect audio, a record audio, a stop recording, and a save recording.
3. The system (102) as claimed in claim 1, wherein the one or more audio data comprises at least one of a speech data, an object sound, a music recording, and a digitally recorded human sounds, wherein the one or more incidents comprises at least one of a gunshot, a scream, a hitting sound, a fire alarm and an explosion.
4. The system (102) as claimed in claim 1, wherein the data cleaning process comprises at least one of a sampling down the one or more audio data to at least 16000 Hz mono-channel audio data from an original frequency.
5. The system (102) as claimed in claim 1, wherein the training process comprises train the mono-channel audio data using the one or more deep-learning model and store the trained the mono-channel audio data.
6. The system (102) as claimed in claim 1, wherein the one or more deep learning model comprises at least one of a long-short term memory (LSTP), a 115 dimensional convolutional neural network (1D CNN), and a 2-dimensional convolutional neural network (2D CNN).
7. The system (102) as claimed in claim 1, wherein the pre-defined format comprises at least one of a wave form file format, an audio interchange file format, an audio file format, and a pulse code modulation file format.
8. The system (102) as claimed in claim 1, wherein the registered contact list comprises at least one of a police station, a pink police patrol, and a cyber cell.
9. The system (102) as claimed in claim1, wherein the communication platform comprises at least one of an e-mail, a Short Message Service (SMS) and a WhatsApp.
10. A method for threat detection using real time audio recording, the method comprising: receiving, by the system (102), at least one data packet from one or more computing device associated with one or more users (110), wherein the at least one data packet pertains to one or more audio data captured 30 from an audio recording module; performing, by the system (102), a data-cleaning process for generating a mono-channel audio data based on the one or more audio data, and execute a training process on the one or more audio data to obtain a trained mono-channel audio data; recording, by the system (102), the one or more audio data in real time by an audio detection module to identify one or more incidents associated with the one or more user, and store the one or more audio data in a pre-defined format; identifying, by the system (102), one or more class associated with the one or more audio data by a prediction module, and detect the one or more audio data as a suspicious event by one or more deep learning model selected by the one or more user based on the trained mono-channel audio data; and generating and transmitting, by the system (102), an alert message to one or more registered contact list through one or more communication platforms based on detection of suspicious event associated with the one or more audio data.