Acoustic event detection method and device, electronic equipment and storage medium
The Kalman filtering and graph neural network model are fused with multimodal sensor data and audio data, combined with CNN and LSTM models to extract features, and match event types in the fingerprint library, solving the problem of poor acoustic event detection in complex scenarios and achieving efficient event detection.
Patent Information
- Application Number
- CN202510524171.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-24
- Publication Date
- 2025-07-22
AI Technical Summary
The existing acoustic event detection technology has poor results in complex scenarios and faces the problems of noise interference, data dependence, inefficiency of multimodal fusion and insufficient real-time performance.
Kalman filtering and graph neural network model are used to perform spatiotemporal fusion of multimodal sensor data and audio data, feature extraction is combined with CNN and LSTM models, and event types are matched in the acoustic event fingerprint library through fingerprint encoding.
It improves the acoustic event detection effect in complex scenarios, improves computing efficiency by reducing dimensional disasters, and supports dynamic update of fingerprint libraries to adapt to new event types.
Smart Images

Figure CN120356484A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio processing, and in particular, to an acoustic event detection method and device, an electronic device, and a storage medium. Background Art
[0002] Acoustic event detection refers to the technology of automatically identifying and classifying specific sound events by analyzing audio signals, such as identifying the sound of broken glass, the cry of a baby, abnormal noises of equipment failures, etc. Its core goal is to locate the start and end times of the target sound in the continuous audio stream and determine its category.
[0003] Currently, the methods of acoustic event detection are mainly divided into traditional methods and deep learning methods. Traditional methods include feature extraction and classifiers, such as MFCC plus SVM or HMM. Deep learning methods are models such as CNN and RNN, and some hybrid models such as CRNN may also be mentioned. However, although these acoustic event detection technologies perform well in simple scenarios, they still face core challenges such as noise interference, data dependence, inefficient multimodal fusion, and insufficient real-time performance, resulting in poor detection effects in complex scenarios. Summary of the Invention
[0004] Embodiments of the present invention provide an acoustic event detection method and device, an electronic device, and a storage medium to solve the problem of poor acoustic event detection effects in complex scenarios.
[0005] In a first aspect, an embodiment of the present invention provides an acoustic event detection method, including: Obtaining multimodal sensor data and audio data representing acoustic events in a target area; Performing spatio-temporal fusion on the multimodal sensor data and the audio data based on a Kalman filter and a graph neural network model to obtain fusion data of the target area; Performing feature extraction on the fusion data based on a trained CNN model and an LSTM model to obtain fusion features representing acoustic events; Performing fingerprint encoding on the fusion features and performing matching in a preset acoustic event fingerprint library based on the fingerprint encoding to obtain the acoustic event type of the target area.
[0006] In a possible implementation manner, obtaining multimodal sensor data and audio data representing acoustic events in a target area includes: Obtaining multimodal sensor data and original audio data in a target area; Identifying the scene type of the target area based on the multimodal sensor data; Filtering the original audio data based on the threshold corresponding to the scene type to obtain audio data representing acoustic events.
[0007] In a possible implementation, based on the Kalman filter and the graph neural network model, spatio-temporal fusion is performed on the multi-modal sensor data and the audio data to obtain the fusion data of the target area, including: Perform state estimation on the multi-modal sensor data based on the Kalman filter to obtain the filtered sensor data; Based on the graph neural network model corresponding to the scene type of the target area, map the filtered sensor data and the audio data to graph nodes to obtain the fusion data of the target area.
[0008] In a possible implementation, before mapping the filtered sensor data and the audio data to graph nodes based on the graph neural network model corresponding to the scene type of the target area to obtain the fusion data of the target area, it further includes: Construct an initial graph neural network based on the spatial relationship of the sensors; For each scene type, train the initial graph neural network based on multiple groups of multi-modal sensor data and audio data representing acoustic events under this scene type to obtain the graph neural network corresponding to this scene type.
[0009] In a possible implementation, based on the trained CNN model and LSTM model, feature extraction is performed on the fusion data to obtain the fusion features representing acoustic events, including: Extract the spatial features of the fusion data based on the CNN model; Extract the temporal features of the fusion data based on the LSTM model; Perform weighted fusion on the spatial features and the temporal features to obtain the fusion features representing acoustic events.
[0010] In a possible implementation, before performing feature extraction on the fusion data based on the trained CNN model and LSTM model to obtain the fusion features representing acoustic events, it further includes: Obtain multiple groups of multi-modal sensor data and audio data representing acoustic events; Train the GAN model based on multiple groups of multi-modal sensor data and audio data representing acoustic events to obtain the trained GAN model; Generate adversarial samples based on the trained GAN model and form a training sample set with multiple groups of multi-modal sensor data and audio data representing acoustic events; Perform adversarial training on the initial CNN model and LSTM model based on the training sample set to obtain the trained CNN model and LSTM model.
[0011] In a possible implementation, there are multiple IoT devices in the target area; fingerprint encoding is performed on the fusion features, and matching is performed in a preset acoustic event fingerprint database based on the fingerprint encoding to obtain the acoustic event type of the target area, including: Each IoT device in the target area respectively performs fingerprint encoding on its own fusion features, and performs matching in a preset acoustic event fingerprint database based on the fingerprint encoding to obtain the acoustic event type detected by the IoT device; The decision tree algorithm is used to fuse the acoustic event types detected by each IoT device to obtain the acoustic event type of the target area.
[0012] In a second aspect, an embodiment of the present invention provides an acoustic event detection device, including: An acquisition module, configured to acquire multimodal sensor data and audio data characterizing acoustic events in the target area; A fusion module, configured to perform spatio-temporal fusion on the multimodal sensor data and the audio data based on a Kalman filter and a graph neural network model to obtain fusion data of the target area; An extraction module, configured to perform feature extraction on the fusion data based on a trained CNN model and an LSTM model to obtain fusion features characterizing acoustic events; A matching module, configured to perform fingerprint encoding on the fusion features, and perform matching in a preset acoustic event fingerprint database based on the fingerprint encoding to obtain the acoustic event type of the target area.
[0013] In a third aspect, an embodiment of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where when the processor executes the computer program, the steps of the method described in the first aspect or any possible implementation manner of the first aspect above are implemented.
[0014] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, where when the computer program is executed by a processor, the steps of the method described in the first aspect or any possible implementation manner of the first aspect above are implemented.
[0015] An embodiment of the present invention provides an acoustic event detection method, apparatus, electronic device, and storage medium. The method realizes the fusion of sensor data and audio data through Kalman filtering and a graph neural network model, and extracts highly discriminative fusion features from the fusion data through a CNN model and an LSTM model, avoiding the dimensionality disaster caused by the splicing of original data. The high-dimensional fusion features are encoded into compact fingerprints, which can quickly match event types in the fingerprint database, improving the calculation efficiency. Moreover, the fingerprint database supports dynamic update, and new event types can be adapted without retraining the model, thereby improving the acoustic event detection effect in complex scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0017] Figure 1 is the implementation flowchart of the acoustic event detection method provided by the embodiment of the present invention; Figure 2 is the structural schematic diagram of the acoustic event detection apparatus provided by the embodiment of the present invention; Figure 3 is the schematic diagram of the electronic device provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.
[0019] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will be described through specific embodiments with reference to the drawings.
[0020] See Figure 1 , which shows the implementation flowchart of the acoustic event detection method provided by the embodiment of the present invention, and is described in detail as follows: Step 101, obtain multi-modal sensor data and audio data representing acoustic events in the target area.
[0021] In this embodiment, multi-modal sensor data refers to heterogeneous data from different types of sensors, such as temperature sensors (ambient temperature), acceleration sensors (vibration intensity), light sensors (light intensity), infrared sensors (human presence detection), etc.
[0022] The audio data characterizing the acoustic event is a preprocessed audio signal that can reflect the acoustic characteristics of the target event, such as the time-frequency characteristics extracted by Mel-Frequency Cepstral Coefficients (MFCC) and Mel-Spectrogram.
[0023] Sensor data provides the physical environment state, and audio data directly captures sound events. The two can cover the physical and acoustic dimensions of the event, providing a basis for subsequent fusion and event detection.
[0024] Step 102: Based on the Kalman filter and the graph neural network model, perform spatio-temporal fusion on the multi-modal sensor data and audio data to obtain the fused data of the target area.
[0025] In this embodiment, the Kalman filter can predict the future state based on the historical data of the sensor, and can eliminate sensor noise. The graph neural network model can map the sensor data and audio data into graph nodes, define the edge weights through physical topology (such as sensor position) or statistical correlation (covariance), establish the spatial dependence between the sensors and the audio, and enhance the event context awareness.
[0026] Step 103: Based on the trained CNN model and LSTM model, extract features from the fused data to obtain the fused features characterizing the acoustic event.
[0027] In this embodiment, the CNN model extracts local spatial features through convolutional kernels and is suitable for processing grid data such as time-frequency graphs. The LSTM model captures the long-term dependence of sequential data through gating mechanisms. The spatial features output by the CNN and the sequential features output by the LSTM can be concatenated or weighted and fused. The combined features cover the spatio-temporal dimensions and reduce the subsequent matching complexity.
[0028] Step 104: Perform fingerprint encoding on the fused features and perform matching in a preset acoustic event fingerprint library based on the fingerprint encoding to obtain the acoustic event type of the target area.
[0029] In this embodiment, fingerprint encoding can map high-dimensional features into low-dimensional compact representations (such as binary hash codes), and common encoding methods include Locality-Sensitive Hashing (LSH).
[0030] During fingerprint matching, calculate the Hamming distance (the number of different bits in the binary code) or cosine similarity between the current fingerprint and the fingerprints in the fingerprint library, and select the closest event type to identify the acoustic event type of the target area.
[0031] In the embodiments of the present invention, the fusion of sensor data and audio data is achieved through Kalman filtering and a graph neural network model, and high-discriminative fusion features are extracted from the fusion data through a CNN model and an LSTM model, avoiding the dimensionality disaster caused by the splicing of raw data. The high-dimensional fusion features are encoded into compact fingerprints, which can quickly match event types in the fingerprint database, improving the calculation efficiency. Moreover, the fingerprint database supports dynamic updates, and new event types can be adapted without retraining the model, thereby improving the acoustic event detection effect in complex scenarios.
[0032] In a possible implementation manner, obtaining multi-modal sensor data and audio data representing acoustic events in a target area includes: Obtaining multi-modal sensor data and raw audio data in the target area; Identifying the scene type of the target area based on the multi-modal sensor data; Filtering the raw audio data based on the threshold corresponding to the scene type to obtain audio data representing acoustic events.
[0033] In this embodiment, environmental parameters can be continuously and real-time collected through sensors such as temperature, humidity, acceleration, and light in the target area to form time-series data (such as sampling once per second), which serves as multi-modal sensor data. At the same time, a microphone array is used to continuously collect raw audio waveforms, which serves as raw audio data. The sensor and audio data need to be synchronized through a unified timestamp or a sliding window.
[0034] The scene type may include home, factory, office, etc. Since the raw audio data has not been denoised and intercepted, background noise in the scene (such as wind noise, human voices) will mask the characteristics of the target acoustic event, resulting in false detection or missed detection. For example, the air conditioner noise in a home environment may be misjudged as equipment failure noise. Common scene types and characteristics are shown in Table 1.
[0035] Table 1
[0036] Based on the sensor characteristics of each scene type, the scene type of the target area can be identified using multi-modal sensor data. For example, the sensor data is input into a scene classifier (such as an SVM or a lightweight CNN) to determine the current scene type. Thus, appropriate methods can be adopted to denoise and intercept the raw audio data, improving the accuracy of acoustic event detection.
[0037] Different scenarios have different background noise characteristics and may require different thresholds to process audio data. Specifically, the threshold can be set according to the typical noise level or event characteristics of the scenario. For example, in a factory environment, the background noise is relatively high, and a higher threshold may be required to avoid false triggers; while in a home environment, the threshold may be lower to detect more subtle sound events. The threshold setting strategies for various scenario types are shown in Table 2.
[0038] Table 2
[0039] In a possible implementation, based on the Kalman filter and the graph neural network model, spatio-temporal fusion is performed on the multi-modal sensor data and audio data to obtain the fusion data of the target area, including: Performing state estimation on the multi-modal sensor data based on the Kalman filter to obtain the filtered sensor data; Based on the graph neural network model corresponding to the scenario type of the target area, mapping the filtered sensor data and audio data into graph nodes to obtain the fusion data of the target area.
[0040] In this embodiment, the sensor data (such as the smoothed sequence of temperature and acceleration sensors) after Kalman filtering can be in the form of [time step, number of sensors, feature dimension]. For example, the temperature data sampled every 0.1 second within 10 seconds is represented as: [100, 5, 1] (5 temperature sensors, single feature).
[0041] The audio data can be a time-frequency graph (such as MFCC or Mel spectrogram), in the form of [time step, frequency channel]. For example, the MFCC after 10-second audio framing can be represented as: [100, 13] (13-dimensional MFCC coefficients).
[0042] Then, each sensor is regarded as a node in the graph, and the node feature is the filtered time-series data (sensor position encoding can be added). The audio data is regarded as a "virtual node" and added to the graph. The node feature is the local statistic (such as mean, variance) of the time-frequency graph, and a strong connection is established between the audio node and the key sensor (such as the motion sensor near the microphone) through the attention weight.
[0043] Through the graph neural network model, the time-frequency features of the audio and the sensor time-series features can be fused at the same semantic level, solving the problem of multi-modal data heterogeneity and realizing acoustic event detection using multi-modal data.
[0044] In a possible implementation, before mapping the filtered sensor data and audio data into graph nodes based on the graph neural network model corresponding to the scenario type of the target area to obtain the fusion data of the target area, it further includes: Construct an initial graph neural network based on the spatial relationship of sensors; For each scenario type, train the initial graph neural network based on multiple sets of multimodal sensor data and audio data representing acoustic events under this scenario type to obtain the graph neural network corresponding to this scenario type.
[0045] In this embodiment, the positions, types, or data characteristics of sensor nodes may be different under different scenarios. Therefore, the structure of the graph (such as the adjacency matrix) or the fusion strategy may need to be dynamically adjusted. For example, in a home scenario, door and window sensors may be more relevant to audio nodes, while in a factory scenario, vibration sensors may be more important.
[0046] Specifically, the initial graph neural network needs to be constructed based on the spatial relationship of sensors to make the graph structure match the actual sensor layout. The connection strength between nodes in the graph needs to match the interaction and influence relationship of each sensor in the scenario. Therefore, for different scenarios, by training the GNN model with corresponding data, it can learn the interaction patterns unique to the scenario, such as the strong correlation between vibration sensors and equipment noise in a factory.
[0047] For example, the steps of constructing the initial graph neural network may include: Node definition: Each sensor is a node, and the feature is the filtered time-series data (such as temperature mean, acceleration variance). The audio data is a virtual node, and the feature is the time-frequency graph statistic (such as spectral centroid, energy entropy).
[0048] Edge definition: Calculate the adjacency matrix based on the physical distance of sensors, Pearson correlation coefficient, or mutual information (such as the closer the distance, the higher the weight).
[0049] The steps of scene-adaptive GNN training may include: Scene classification data preparation: Collect multimodal datasets for different scenarios (homes, factories, etc.) and label the scenario types.
[0050] Model training: For each scenario, use the corresponding data to train the GNN model (such as GAT or GCN), and the loss function is the classification cross-entropy.
[0051] Input: Filtered sensor data + audio features.
[0052] Output: Scenario type label (such as "home", "factory").
[0053] Train a dedicated GNN model for each scenario and store it as pre-trained weights.
[0054] Through scene - adaptive training, high robustness, strong scene adaptability, and high efficiency of acoustic event detection can be achieved, which is applicable to complex environments such as smart homes and industrial monitoring.
[0055] In a possible implementation, based on the trained CNN model and LSTM model, feature extraction is performed on the fusion data to obtain fusion features representing acoustic events, including: Extracting the spatial features of the fusion data based on the CNN model; Extracting the temporal features of the fusion data based on the LSTM model; Weightedly fusing the spatial features and the temporal features to obtain fusion features representing acoustic events.
[0056] In this embodiment, the CNN model and the LSTM model have feature complementarity. The CNN is suitable for capturing local details of acoustic events, while the LSTM can model the temporal evolution of the environmental state.
[0057] Specifically, the CNN model inputs the audio time - frequency map and outputs a flattened feature vector, and the LSTM model inputs the sensor temporal data and outputs the hidden state of the last time step. Then, the spatial features and the temporal features are concatenated into joint features, and weights are dynamically assigned through the attention mechanism, which can achieve spatio - temporal joint representation covering the physical and acoustic dimensions of events, and the information dimension is reduced, reducing the computational cost and improving the matching efficiency.
[0058] In a possible implementation, before performing feature extraction on the fusion data based on the trained CNN model and LSTM model to obtain fusion features representing acoustic events, it further includes: Obtaining multiple groups of multi - modal sensor data and audio data representing acoustic events; Training the GAN model based on multiple groups of multi - modal sensor data and audio data representing acoustic events to obtain a trained GAN model; Generating adversarial samples based on the trained GAN model and forming a training sample set with multiple groups of multi - modal sensor data and audio data representing acoustic events; Performing adversarial training on the initial CNN model and LSTM model based on the training sample set to obtain the trained CNN model and LSTM model.
[0059] In this embodiment, the GAN model (Generative Adversarial Network) is an adversarial framework composed of a generator (Generator) and a discriminator (Discriminator), where the generator generates realistic data and the discriminator distinguishes between real and generated data.
[0060] Using multiple sets of multi-modal sensor data and audio data representing acoustic events, alternately optimize the generator (to deceive the discriminator) and the discriminator (to distinguish between true and false) until convergence to obtain a trained GAN model. During the training process, the distribution difference (such as the Wasserstein distance) between the generated data and the real data can be minimized.
[0061] Adversarial samples are data generated by GANs. Their statistical characteristics are similar to those of real data but contain perturbations, and are used to simulate complex scenarios. Thus, in the case of insufficient sample numbers, a diverse audio set can be generated to improve the training efficiency of CNN models and LSTM models.
[0062] When training CNN models and LSTM models, input a mixed training set (real data + adversarial samples), jointly optimize the CNN model and the LSTM model, and use the cross-entropy loss and the adversarial loss as the loss function together.
[0063] In a possible implementation, there are multiple IoT devices in the target area; fingerprint code the fusion features, and perform matching in a preset acoustic event fingerprint library based on the fingerprint code to obtain the acoustic event types in the target area, including: Each IoT device in the target area respectively fingerprint codes its own fusion features, and performs matching in a preset acoustic event fingerprint library based on the fingerprint code to obtain the acoustic event type detected by this IoT device; Use the decision tree algorithm to fuse the acoustic event types detected by each IoT device to obtain the acoustic event type in the target area.
[0064] In this embodiment, multiple IoT devices may be deployed in the target area, and each device can independently perform acoustic event detection to obtain its own results. Integrating these results can improve accuracy and robustness.
[0065] When using the decision tree algorithm to fuse the acoustic event types detected by each IoT device, it is necessary to collect the detection results of each device and their context information, such as the location of the device, the timestamp, the sensor data, etc. The location information of the device can be transformed into features of distance or area division. For example, the closer the devices are to each other, the more relevant their results may be. In addition, time synchronization is also a key point to ensure that the data of all devices are aligned in time and avoid misjudgment caused by time differences.
[0066] In a specific embodiment, the data input to the decision tree may include the acoustic event types (classification labels), confidence levels (probability values), and context features output by each IoT device.
[0067] The context features may include the following: Device location (such as coordinates or area division); Device historical accuracy rate (e.g., the accuracy rate of D1 in detecting "broken glass" in the past was 90%); Environmental parameters (such as the current environmental noise level, temperature); Timestamp (used to synchronize multi-device data).
[0068] The decision tree generated from the training data contains the following typical rules: Rule 1: If the number of votes for "broken glass" ≥ 2 and the average confidence level > 0.8, then it is determined as "broken glass".
[0069] Rule 2: If the number of votes for "broken glass" = 1, but the historical accuracy rate of the device with the highest confidence level > 85%, then it is determined as "broken glass".
[0070] Rule 3: If there is a conflict in the device detection results (e.g., the number of votes for "broken glass" and "background noise" is the same) and the event source is close to a certain device, then adopt the result of the device with a shorter distance.
[0071] Finally, the decision tree outputs the final acoustic event type.
[0072] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0073] The following is an embodiment of the device of the present invention. For the details not described in detail, reference can be made to the corresponding method embodiments above.
[0074] Figure 2 The structural schematic diagram of the acoustic event detection device provided by the embodiment of the present invention is shown. For the sake of convenience of description, only the parts related to the embodiment of the present invention are shown and are described in detail as follows: As Figure 2 shown, the acoustic event detection device 2 includes: An acquisition module 21, configured to acquire multi-modal sensor data and audio data characterizing acoustic events in a target area; A fusion module 22, configured to perform spatio-temporal fusion on the multi-modal sensor data and audio data based on a Kalman filter and a graph neural network model to obtain fusion data of the target area; An extraction module 23, configured to perform feature extraction on the fusion data based on a trained CNN model and an LSTM model to obtain fusion features characterizing acoustic events; A matching module 24, configured to perform fingerprint encoding on the fusion features and perform matching in a preset acoustic event fingerprint library based on the fingerprint encoding to obtain the acoustic event type of the target area.
[0075] In a possible implementation, the acquisition module 21 is specifically configured to: Acquire multimodal sensor data and original audio data in the target area; Identify the scene type of the target area based on the multimodal sensor data; Filter the original audio data based on the threshold corresponding to the scene type to obtain audio data representing acoustic events.
[0076] In a possible implementation, the fusion module 22 is specifically configured to: Perform state estimation on the multimodal sensor data based on Kalman filtering to obtain filtered sensor data; Based on the graph neural network model corresponding to the scene type of the target area, map the filtered sensor data and audio data to graph nodes to obtain the fusion data of the target area.
[0077] In a possible implementation, the fusion module 22 is further configured to: Before mapping the filtered sensor data and audio data to graph nodes based on the graph neural network model corresponding to the scene type of the target area to obtain the fusion data of the target area, construct an initial graph neural network based on the spatial relationship of the sensors; For each scene type, train the initial graph neural network based on multiple sets of multimodal sensor data and audio data representing acoustic events under this scene type to obtain the graph neural network corresponding to this scene type.
[0078] In a possible implementation, based on the trained CNN model and LSTM model, perform feature extraction on the fusion data to obtain fusion features representing acoustic events, where the extraction module 23 is specifically configured to: Extract the spatial features of the fusion data based on the CNN model; Extract the temporal features of the fusion data based on the LSTM model; Perform weighted fusion on the spatial features and temporal features to obtain fusion features representing acoustic events.
[0079] In a possible implementation, the extraction module 23 is further configured to: Before performing feature extraction on the fusion data based on the trained CNN model and LSTM model to obtain fusion features representing acoustic events, acquire multiple sets of multimodal sensor data and audio data representing acoustic events; Train the GAN model based on multiple sets of multimodal sensor data and audio data representing acoustic events to obtain the trained GAN model; Generate adversarial samples based on a trained GAN model, and form a training sample set with multiple groups of multimodal sensor data and audio data representing acoustic events; Perform adversarial training on the initial CNN model and LSTM model based on the training sample set to obtain the trained CNN model and LSTM model.
[0080] In a possible implementation, there are multiple IoT devices in the target area; the matching module 24 is specifically configured to: For each IoT device in the target area, perform fingerprint encoding on its respective fusion feature, and perform matching in a preset acoustic event fingerprint library based on the fingerprint encoding to obtain the type of acoustic event detected by the IoT device; Use the decision tree algorithm to fuse the types of acoustic events detected by each IoT device to obtain the type of acoustic event in the target area.
[0081] In the embodiment of the present invention, the fusion of sensor data and audio data is realized through Kalman filtering and a graph neural network model, and high-discriminative fusion features are extracted from the fusion data through the CNN model and the LSTM model, avoiding the dimensionality disaster caused by the splicing of raw data. The high-dimensional fusion features are encoded into compact fingerprints, which can quickly match the event types in the fingerprint library, improving the calculation efficiency. And the fingerprint library supports dynamic update, and the new event types can be adapted without retraining the model, thereby improving the acoustic event detection effect in complex scenarios.
[0082] Figure 3 It is a schematic diagram of an electronic device provided by an embodiment of the present invention. As Figure 3 shown, the electronic device 3 of this embodiment includes: a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30. When the processor 30 executes the computer program 32, the steps in the above-mentioned embodiments of each acoustic event detection method are implemented. Alternatively, when the processor 30 executes the computer program 32, the functions of each module / unit in the above-mentioned device embodiments are implemented.
[0083] Exemplarily, the computer program 32 can be divided into one or more modules / units, and the one or more modules / units are stored in the memory 31 and executed by the processor 30 to complete the present invention. The one or more modules / units can be a series of computer program instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program 32 in the electronic device 3.
[0084] The electronic device 3 may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The electronic device 3 may include, but is not limited to, a processor 30 and a memory 31. Those skilled in the art can understand that Figure 3 merely examples of the electronic device 3, which do not constitute a limitation on the electronic device 3, may include more or fewer components than shown in the figure, or combine certain components, or different components. For example, the electronic device may further include input / output devices, network access devices, a bus, etc.
[0085] The so-called processor 30 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0086] The memory 31 may be an internal storage unit of the electronic device 3, such as a hard disk or memory of the electronic device 3. The memory 31 may also be an external storage device of the electronic device 3, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 3. Further, the memory 31 may also include both the internal storage unit and the external storage device of the electronic device 3. The memory 31 is used to store the computer program and other programs and data required by the electronic device. The memory 31 may also be used to temporarily store the data that has been output or will be output.
[0087] Those skilled in the art can clearly understand that, for the convenience and conciseness of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of each functional unit and module are only for the convenience of mutual distinction and do not limit the protection scope of this application. The specific working processes of the units and modules in the above system can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.
[0088] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0089] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in the form of hardware or software depends on the specific application and design constraints of the technical solution. Professionals can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0090] In the embodiments provided by the present invention, it should be understood that the disclosed device / electronic device and method can be implemented in other ways. For example, the device / electronic device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the device or unit can be in an electrical, mechanical or other form.
[0091] The unit described as a separated component may or may not be physically separated, and the component displayed as a unit may or may not be a physical unit, that is, it may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0092] In addition, in each embodiment of the present invention, each functional unit can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0093] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, to implement all or part of the processes in the above-described embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-described various embodiments of the acoustic event detection method can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device that can carry the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0094] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. An acoustic event detection method, characterized in that, Including: Obtain multimodal sensor data and audio data representing acoustic events in the target area; Based on the Kalman filter and the graph neural network model, perform spatio-temporal fusion on the multimodal sensor data and the audio data to obtain the fusion data of the target area; Based on the trained CNN model and LSTM model, extract features from the fusion data to obtain fusion features representing acoustic events; Perform fingerprint coding on the fusion features and perform matching in a preset acoustic event fingerprint library based on the fingerprint coding to obtain the acoustic event type of the target area.
2. The acoustic event detection method according to claim 1, wherein The obtaining multimodal sensor data and audio data representing acoustic events in the target area includes: Obtain multimodal sensor data and original audio data in the target area; Identify the scene type of the target area based on the multimodal sensor data; Filter the original audio data based on the threshold corresponding to the scene type to obtain audio data representing acoustic events.
3. The acoustic event detection method according to claim 2, wherein The performing spatio-temporal fusion on the multimodal sensor data and the audio data based on the Kalman filter and the graph neural network model to obtain the fusion data of the target area includes: Perform state estimation on the multimodal sensor data based on the Kalman filter to obtain filtered sensor data; Based on the graph neural network model corresponding to the scene type of the target area, map the filtered sensor data and the audio data to graph nodes to obtain the fusion data of the target area.
4. The acoustic event detection method according to claim 3, wherein Before the mapping the filtered sensor data and the audio data to graph nodes based on the graph neural network model corresponding to the scene type of the target area to obtain the fusion data of the target area, it further includes: Construct an initial graph neural network based on the spatial relationship of the sensors; For each scene type, train the initial graph neural network based on multiple sets of multimodal sensor data and audio data representing acoustic events under this scene type to obtain the graph neural network corresponding to this scene type.
5. The acoustic event detection method according to claim 1, wherein The extracting features from the fusion data based on the trained CNN model and LSTM model to obtain fusion features representing acoustic events includes: Extract the spatial features of the fusion data based on the CNN model; Extract the temporal features of the fusion data based on the LSTM model; Perform weighted fusion on the spatial features and the temporal features to obtain fusion features representing acoustic events.
6. The acoustic event detection method according to claim 5, wherein Before the extracting features from the fusion data based on the trained CNN model and LSTM model to obtain fusion features representing acoustic events, it further includes: Obtain multiple sets of multimodal sensor data and audio data representing acoustic events; Train the GAN model based on the multiple sets of multimodal sensor data and audio data representing acoustic events to obtain a trained GAN model; Generate adversarial samples based on the trained GAN model and form a training sample set with the multiple sets of multimodal sensor data and audio data representing acoustic events. Perform adversarial training on the initial CNN model and LSTM model based on the training sample set to obtain the trained CNN model and LSTM model.
7. The acoustic event detection method according to claim 1, characterized in that, There are multiple IoT devices in the target area; the fingerprint encoding of the fusion feature and the matching in a preset acoustic event fingerprint library based on the fingerprint encoding to obtain the acoustic event type in the target area include: Each IoT device in the target area respectively performs fingerprint encoding on its own fusion feature and performs matching in a preset acoustic event fingerprint library based on the fingerprint encoding to obtain the acoustic event type detected by the IoT device. Use the decision tree algorithm to fuse the acoustic event types detected by each IoT device to obtain the acoustic event type in the target area.
8. An acoustic event detection device, characterized in that, Include: An acquisition module for acquiring multi-modal sensor data in the target area and audio data representing acoustic events. A fusion module for performing spatio-temporal fusion on the multi-modal sensor data and the audio data based on the Kalman filter and the graph neural network model to obtain the fusion data in the target area. An extraction module for extracting features from the fusion data based on the trained CNN model and LSTM model to obtain the fusion feature representing the acoustic event. A matching module for performing fingerprint encoding on the fusion feature and performing matching in a preset acoustic event fingerprint library based on the fingerprint encoding to obtain the acoustic event type in the target area.
9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7 above.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7 above.
Citation Information
Cited By
Equipment fault detection method and device
CN121438858A