Acoustic event localization method, apparatus, system, and computer-readable storage medium
By acquiring multimodal fusion data through high-order surround acoustic spherical array nodes and using a deep learning model with multi-head attention mechanism for feature extraction and recognition, the problem of inaccurate localization and high computational complexity in complex environments in existing acoustic localization technologies is solved, achieving high-precision and low-cost acoustic event localization.
Patent Information
- Application Number
- CN202510504831.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-22
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2045-04-22
AI Technical Summary
Existing acoustic localization technologies struggle to accurately locate acoustic events in complex environments, especially when multiple events occur concurrently, signal temporal aliasing leads to ambiguous solutions, and distributed iterative algorithms increase computational complexity and affect real-time performance.
A high-order surround acoustic spherical array node is used to acquire multimodal fusion data. The localization model is combined with the feature extraction module and the spatiotemporal fusion module. A deep learning model with multi-head attention mechanism is used for feature extraction and recognition, and the type and coordinates of the acoustic event are output.
It significantly improves the accuracy and real-time performance of acoustic event localization in complex environments, simplifies the calculation process, reduces the transmission burden, and improves the system's cost-effectiveness.
Smart Images

Figure CN120370260B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of acoustic localization technology, and in particular to an acoustic event localization method, apparatus, system, and computer-readable storage medium. Background Technology
[0002] Acoustic event localization technology accurately determines the location of an event by analyzing the characteristics and propagation patterns of sound signals, and has a wide range of applications.
[0003] Some existing acoustic localization technologies are prone to aliasing in the signal time domain, leading to ambiguous solutions in the event correlation matrix and making it difficult to output accurate time localization. Other acoustic localization technologies construct sound field statistical models based on Gaussian mixture models. This model assumes that the spatial distribution of sound sources is separable, which limits its application in complex scenarios. Still other acoustic localization technologies introduce distributed iterative algorithms, which increases computational complexity and affects real-time performance. Summary of the Invention
[0004] The purpose of this invention is to provide an acoustic event localization method, apparatus, system, and computer-readable storage medium, which acquires accurate multimodal fusion data, and uses a localization model with a feature extraction module and a spatiotemporal fusion module for feature extraction and identification, accurately outputs the type and location of the acoustic event, simplifies the calculation, and significantly improves the phenotypic performance of the model in complex environments.
[0005] In a first aspect, the present invention provides an acoustic event localization method, applied to an acoustic event localization system, the acoustic event localization system comprising: a data processing subsystem and multiple high-order surround acoustic spherical array nodes; the method comprising:
[0006] High-order surround acoustic spherical array nodes acquire multimodal fusion data; wherein, the multimodal fusion data includes: acoustic field information of acoustic events, temporal information of acoustic events, position information of high-order surround acoustic spherical array nodes, attitude information of high-order surround acoustic spherical array nodes, and inference type of acoustic events;
[0007] The high-order surround acoustic spherical array nodes compress the multimodal fusion data and send it to the data processing subsystem;
[0008] The data processing subsystem performs feature extraction and recognition on multimodal fusion data based on a pre-trained multi-source data fusion localization model, and outputs the type and coordinates of acoustic events. The multi-source data fusion localization model includes a feature extraction module and a spatiotemporal fusion module. The feature extraction module includes a temporal branch and a spatial branch. The spatiotemporal fusion module is a deep learning model based on a multi-head attention mechanism.
[0009] In some preferred embodiments of the present invention, the high-order surround acoustic spherical array node includes: a spherical microphone array, a real-time dynamic differential system, an attitude sensor, and a sound field information processor; the step of the high-order surround acoustic spherical array node acquiring multimodal fusion data includes:
[0010] Acquiring sound field information of acoustic events based on spherical microphone arrays;
[0011] The position information of the nodes of the high-order surround acoustic spherical array is obtained based on a real-time dynamic differential system.
[0012] Attitude information of high-order surround acoustic spherical array nodes is obtained based on attitude sensors;
[0013] The reasoning type for determining acoustic events is based on the sound field information processor.
[0014] In some preferred embodiments of the present invention, the acoustic event localization system further includes: a meteorological monitoring node; the method further includes:
[0015] Meteorological monitoring nodes acquire meteorological information; the meteorological information includes at least one of the following: wind direction, wind speed, temperature, relative humidity, and air pressure.
[0016] The meteorological monitoring nodes send meteorological information to the data processing subsystem.
[0017] In some preferred embodiments of the present invention, the method further includes:
[0018] The data processing subsystem corrects the sound field information based on meteorological information.
[0019] In some preferred embodiments of the present invention, the high-order surround acoustic spherical array node transmits the compressed multimodal fusion data to the data processing subsystem through low-power wide-area transmission technology.
[0020] In some preferred embodiments of the present invention, the temporal branch captures the temporal correlation characteristics of the multimodal fusion data through a bidirectional gated logic unit; the spatial branch extracts the directional distribution features of the multimodal fusion data using a two-dimensional convolutional kernel.
[0021] In some preferred embodiments of the present invention, the spatiotemporal fusion module splits the input vector into multiple subspaces, then concatenates the results and performs a linear transformation to obtain the output result; wherein the attention patterns of each subspace are different.
[0022] Secondly, the present invention provides an acoustic event localization device for use in an acoustic event localization system, the acoustic event localization system comprising: a data processing subsystem and multiple high-order surround acoustic spherical array nodes; the device comprising:
[0023] The data acquisition module is used to acquire multimodal fusion data from the high-order surround acoustic spherical array nodes. The multimodal fusion data includes: acoustic field information of acoustic events, time information of acoustic events, position information of high-order surround acoustic spherical array nodes, attitude information of high-order surround acoustic spherical array nodes, and inference type of acoustic events.
[0024] The data transmission module is used to compress the multimodal fusion data from the high-order surround acoustic spherical array nodes and send it to the data processing subsystem.
[0025] The data analysis module is used by the data processing subsystem to extract and identify features from multimodal fusion data based on a pre-trained multi-source data fusion localization model, and output the type and coordinates of acoustic events. The multi-source data fusion localization model includes a feature extraction module and a spatiotemporal fusion module. The feature extraction module includes a temporal branch and a spatial branch. The spatiotemporal fusion module is a deep learning model based on a multi-head attention mechanism.
[0026] Thirdly, the present invention provides an acoustic event localization system, including a processor and a memory, wherein the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the acoustic event localization method provided in the first aspect above.
[0027] Fourthly, the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the acoustic event localization method provided in the first aspect.
[0028] This invention brings the following beneficial effects:
[0029] This invention provides an acoustic event localization method, apparatus, system, and computer-readable storage medium, applied to an acoustic event localization system. The acoustic event localization system includes: a data processing subsystem and multiple high-order surround acoustic spherical array nodes; the method includes: the high-order surround acoustic spherical array nodes acquiring multimodal fusion data; wherein, the multimodal fusion data includes: sound field information of the acoustic event, time information of the acoustic event, position information of the high-order surround acoustic spherical array nodes, attitude information of the high-order surround acoustic spherical array nodes, and inference type of the acoustic event; the high-order surround acoustic spherical array nodes compress the multimodal fusion data and then send it to the data processing subsystem. The system's data processing subsystem extracts and identifies features from multimodal fusion data based on a pre-trained multi-source data fusion localization model, outputting the type and coordinates of acoustic events. The multi-source data fusion localization model includes a feature extraction module and a spatiotemporal fusion module. The feature extraction module includes temporal and spatial branches. The spatiotemporal fusion module is a deep learning model based on a multi-head attention mechanism. The system acquires accurate multimodal fusion data and utilizes the localization model with both feature extraction and spatiotemporal fusion modules for feature extraction and identification, accurately outputting the type and location of acoustic events. This simplifies computation and significantly improves the model's phenotypic performance even in complex environments. Attached Figure Description
[0030] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0031] Figure 1 A schematic diagram of the structure of an acoustic event localization system provided in an embodiment of the present invention;
[0032] Figure 2 A flowchart of an acoustic event localization method provided in an embodiment of the present invention;
[0033] Figure 3 A block diagram of a high-order surround acoustic spherical array acoustic sensing system based on embedded AI inference is provided for an embodiment of the present invention.
[0034] Figure 4 A block diagram illustrating the composition of a meteorological node provided in an embodiment of the present invention;
[0035] Figure 5 This is a schematic diagram of the structure of a multi-source data fusion positioning model provided in an embodiment of the present invention;
[0036] Figure 6This is a schematic diagram of the structure of an acoustic event localization device provided in an embodiment of the present invention;
[0037] Figure 7 This is a schematic diagram of the structure of an acoustic event localization system provided in an embodiment of the present invention.
[0038] Icons: 310 - Data acquisition module; 320 - Data transmission module; 330 - Data analysis module; 400 - Memory; 401 - Processor; 402 - Bus; 403 - Communication interface. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.
[0040] Therefore, the following detailed description of the embodiments of the invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely to illustrate selected embodiments of the invention. All other embodiments obtained by those skilled in the art based on the embodiments of the invention without inventive effort are within the scope of protection of the invention.
[0041] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0042] In the description of this invention, it should be noted that the terms "center," "upper," "lower," "left," "right," "vertical," "horizontal," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, or the orientation or positional relationship commonly used when the product of this invention is in use. They are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention. In addition, the terms "first," "second," "third," etc., are only used to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0043] Furthermore, terms such as "horizontal," "vertical," and "sag" do not imply that components must be absolutely horizontal or suspended, but rather that they can be slightly tilted. For example, "horizontal" simply means that its direction is more horizontal relative to "vertical," and does not mean that the structure must be completely horizontal, but can be slightly tilted.
[0044] In the description of this invention, it should also be noted that, unless otherwise explicitly specified and limited, the terms "set," "install," "connect," and "link" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0045] Acoustic event localization technology accurately determines the location of events by analyzing the characteristics and propagation patterns of sound signals, and has a wide range of applications. In security monitoring, it can quickly locate sudden events (such as gunshots and explosions); in industrial inspection, it improves the accuracy of equipment fault diagnosis; and in environmental monitoring and noise control, it contributes to improving environmental quality. Furthermore, acoustic event localization technology offers rapid response and precise positioning, and its widespread application in multiple fields demonstrates significant practical and social value.
[0046] Some existing technologies employ a distributed single-channel microphone network architecture, using asynchronous clock calibration to achieve synchronous signal acquisition from multiple nodes. A time-domain energy detection algorithm is used to independently detect acoustic events at each node and time-stamp them. This system utilizes a cross-node event correlation algorithm to establish a spatiotemporal mapping relationship and combines it with Time Difference of Arrival (TDOA) estimation for sound source localization. This method has advantages such as simple hardware architecture, requiring only a single-channel acquisition unit, and low cost; simultaneously, the event-driven processing mechanism effectively reduces data transmission bandwidth requirements. However, in the case of multiple concurrent events, signal time-domain aliasing may occur, leading to ambiguous solutions in the event correlation matrix. Furthermore, the time delay measurement accuracy of TDOA estimation is limited in non-stationary noise environments, and without establishing a physical model of the sound source, it is difficult to handle the multi-source separation problem in complex acoustic scenarios.
[0047] Some existing technologies construct sound field statistical models based on Gaussian mixture models (GMMs), iteratively estimate the sound source energy distribution using the EM algorithm, and introduce robust mean consensus methods to achieve cooperative parameter estimation in distributed sensor networks. This method improves robustness to non-Gaussian noise through statistical modeling and effectively suppresses the influence of outliers in individual node measurements using consensus protocols. However, the GMM model assumes the separability of the spatial distribution of sound sources, which limits its application in reverberant scenarios. Furthermore, this method does not consider the dispersion characteristics of sound wave propagation, leading to model mismatch issues in broadband sound source localization. In addition, the introduction of distributed iterative algorithms increases computational complexity and affects real-time performance.
[0048] In summary, to address the challenge of multi-source acoustic localization in complex environments, a method for acoustic event localization based on a distributed higher-order ambisonics (HOA) array sensor network is proposed. This method acquires acoustic signals through a distributed acoustic sensor network composed of multiple HOA arrays. Each node possesses high directivity with spatial isotropy, effectively suppressing multipath effects and clutter interference in complex environments. Simultaneously, each node is equipped with a high-precision BeiDou positioning system and a precise timestamp acquisition module, enabling real-time acquisition of the temporal and geographical information of the acoustic signals. An RK3588 embedded AI inference processor is also locally deployed, supporting functions such as event type recognition, sound field distribution estimation, and event occurrence time inference. Finally, the acoustic event type, location, geographical, and time information detected by each node is transmitted to the network aggregation center node. Through multi-source data fusion based on a multimodal neural network model, the multi-source time correlation problem is effectively solved, significantly improving the accuracy of acoustic event localization. This method has high practical value and application prospects for sound source localization in complex environments.
[0049] The following detailed description of some embodiments of the present invention is provided in conjunction with the accompanying drawings. Unless otherwise specified, the following embodiments and features can be combined with each other.
[0050] Example 1
[0051] This invention provides an acoustic event localization method, applied to an acoustic event localization system. See [link to relevant documentation]. Figure 1 The diagram shown is a structural schematic of an acoustic event localization system provided by an embodiment of the present invention. The acoustic event localization system includes: a data processing subsystem and multiple high-order surround acoustic spherical array nodes; the data processing subsystem is used for multi-source information fusion processing.
[0052] See Figure 2 The flowchart shown in this embodiment of the invention provides an acoustic event localization method, which includes:
[0053] Step S102: The high-order surround acoustic spherical array node acquires multimodal fusion data; wherein, the multimodal fusion data includes: acoustic field information of acoustic events, time information of acoustic events, position information of high-order surround acoustic spherical array node, attitude information of high-order surround acoustic spherical array node, and inference type of acoustic events.
[0054] Specifically, multimodal fusion data can be collected in real time or historical data can be obtained.
[0055] Furthermore, in some preferred embodiments of the present invention, the high-order surround acoustic spherical array node includes: a spherical microphone array, a real-time dynamic differential system, an attitude sensor, and a sound field information processor; the step of the high-order surround acoustic spherical array node acquiring multimodal fusion data includes: acquiring sound field information of acoustic events based on the spherical microphone array; acquiring position information of the high-order surround acoustic spherical array node based on the real-time dynamic differential system; acquiring attitude information of the high-order surround acoustic spherical array node based on the attitude sensor; and determining the inference type of the acoustic events based on the sound field information processor.
[0056] Specifically, each high-order surround acoustic spherical array node integrates a spherical microphone array, a BeiDou real-time dynamic differential (RTK) system, an electronic compass, an attitude sensor, and a low-power wide-area transmission unit. These high-order surround acoustic spherical array nodes are evenly deployed around the monitoring area, forming a coverage network.
[0057] The high-order surround acoustic spherical array nodes have uniform directivity, which significantly improves the accuracy of sound source localization and the ability to resist noise / multipath interference; the high-precision distributed synchronous acquisition system based on Beidou achieves high-precision time synchronization and positioning.
[0058] Furthermore, this embodiment of the invention provides a high-order surround acoustic spherical array acoustic sensing system based on embedded AI inference, for processing data locally, see [link to relevant documentation]. Figure 3The illustrated embodiment of the present invention provides a block diagram of a high-order surround sound spherical array acoustic sensing system based on embedded AI inference. This system employs a high-order Ambisonics (HOA) acoustic sensing architecture and constructs a three-dimensional sound field monitoring platform by integrating novel hardware modules and intelligent signal processing algorithms. The core sensing unit consists of a spherical array of 32 omnidirectional MEMS microphones, enabling fifth-order Ambisonics ("panoramic sound technology" or "omnidirectional sound technology," an audio technology for recording, transmitting, and playing back three-dimensional spatial sound information) sound field reconstruction, and is capable of 360° spatial sound field sampling. The multimodal data fusion acquisition system integrates a BeiDou-3 RTK positioning module (positioning accuracy ±1cm) and a Micro-Electro-Mechanical System (MEMS) inertial measurement unit (including a three-axis electronic compass and a six-degree-of-freedom attitude sensor). Simultaneously, the system uses a Xilinx Artix-7 FPGA to achieve nanosecond-level synchronous triggering of multi-channel ADCs and implements 24-bit / 192kHz audio digitization processing through a multi-channel codec. The system also features a Rockchip RK3588 NPU chip, supporting edge computing processing, and includes a dual-core Cortex-A76 processor and a 6 TOPS NPU. The system can locally perform real-time sound source detection and classification based on YOLO-Sound, as well as multi-channel beamforming and Ambisonics B-Format encoding.
[0059] For further details, please refer to [link / reference]. Figure 1 In some preferred embodiments of the present invention, the acoustic event localization system further includes: a meteorological monitoring node; the method further includes: the meteorological monitoring node acquiring meteorological information; wherein the meteorological information includes at least one of the following: wind direction, wind speed, temperature, relative humidity, and air pressure; the meteorological monitoring node sends the meteorological information to the data processing subsystem.
[0060] Specifically, the system also includes meteorological monitoring nodes, communication facilities, power supply equipment, and a data aggregation and processing center. The meteorological monitoring nodes integrate multiple sensors such as thermometers, hygrometers, anemometers, and barometers; the communication system adopts a combination of long-distance low-power wide-area transmission (LoRa) communication and signal relay nodes; the power supply system is powered by batteries and solar power generation devices to ensure long-term stable operation of the system; and the data aggregation and processing center is responsible for data fusion and processing.
[0061] Specifically, meteorological monitoring nodes aim to achieve comprehensive, real-time monitoring of key meteorological elements such as wind direction, wind speed, temperature, relative humidity, and air pressure. (See also...) Figure 4The illustrated embodiment of the present invention provides a block diagram of a meteorological node. This node employs an advanced ARM processor to efficiently integrate data from multiple meteorological sensors and simultaneously acquire precise time and geographic location information provided by the BeiDou Navigation Satellite System. All data is stably transmitted to the central control node via LoRa communication technology. The meteorological monitoring subsystem consists of a series of precision components, including professional meteorological sensors, a miniature meteorological data acquisition unit, a stable power supply system, a lightweight Stevenson screen, a weather-resistant field protective box, and a robust stainless steel support frame. Among these, wind speed and wind direction sensors are precision devices specifically designed for the meteorological field.
[0062] Furthermore, in some preferred embodiments of the present invention, the method further includes: a data processing subsystem correcting the sound field information based on meteorological information.
[0063] Specifically, all collected meteorological data are based on the calculation of sound speed and Doppler distortion, providing an important reference for subsequent compensation and correction of the sound propagation model.
[0064] In step S104, the high-order surround acoustic spherical array node compresses the multimodal fusion data and sends it to the data processing subsystem.
[0065] Specifically, the system uses data compression technology to compress the sound event type, occurrence time, and sound field information, effectively reducing the transmission burden, lowering system costs, and improving deployment flexibility.
[0066] Furthermore, in some preferred embodiments of the present invention, the high-order surround acoustic spherical array node transmits the compressed multimodal fusion data to the data processing subsystem through low-power wide-area transmission technology.
[0067] Specifically, the low-power wide-area transmission system uses the SX1262 LoRa chipset, supporting LDPC forward error correction and frequency hopping spread spectrum technology. The data transmission frame structure includes: BeiDou spatiotemporal reference (UTC timestamp, WGS84 coordinates), node attitude parameters (pitch / roll / yaw angles), and acoustic feature vectors (acoustic fingerprint information, sound field spatial information, event type). This embodiment combines spatial sound field perception technology with high-precision spatiotemporal information, and achieves edge extraction of acoustic features through embedded processing. Compared with traditional solutions, it greatly reduces the amount of raw data transmission while maintaining complete spatial sound field reconstruction capabilities, providing a cost-effective positioning node solution for wide-area acoustic event monitoring networks.
[0068] Step S106: The data processing subsystem performs feature extraction and recognition on the multimodal fusion data based on the pre-trained multi-source data fusion localization model, and outputs the type and coordinates of the acoustic event; wherein, the multi-source data fusion localization model includes: a feature extraction module and a spatiotemporal fusion module; the feature extraction module includes: a temporal branch and a spatial branch; the spatiotemporal fusion module is a deep learning model based on a multi-head attention mechanism.
[0069] For details, see Figure 5 The diagram shown is a structural schematic of a multi-source data fusion localization model provided by an embodiment of the present invention. The multi-source data fusion localization model based on a multimodal neural network model has a core architecture consisting of a feature extraction module and a spatiotemporal fusion module.
[0070] Furthermore, in some preferred embodiments of the present invention, the temporal branch captures the temporal correlation characteristics of the multimodal fusion data through a bidirectional gated logic unit; the spatial branch extracts the directional distribution features of the multimodal fusion data using a two-dimensional convolutional kernel.
[0071] Furthermore, in some preferred embodiments of the present invention, the spatiotemporal fusion module splits the input vector into multiple subspaces, then concatenates the results and performs a linear transformation to obtain the output result; wherein, the attention patterns of each subspace are different.
[0072] Specifically, firstly, multimodal data collected by distributed sensing nodes is aggregated through a LoRa wireless communication network, including acoustic event type features, high-precision timestamp sequences, sound field information features, and geographic information metadata. To address the feature alignment problem of heterogeneous data, a dual-stream feature extractor is constructed using a Convolutional Recurrent Neural Network (CRNN): the temporal branch captures the temporal correlation characteristics of the event sequence through a bidirectional GRU, while the spatial branch extracts the directional distribution features of the acoustic signal using a two-dimensional convolutional kernel. Finally, a feature projection layer maps the multi-source data to a unified 128-dimensional latent space. In the feature fusion stage, a Transformer architecture based on a multi-head attention mechanism is innovatively introduced. The self-attention mechanism dynamically captures long-distance dependencies by calculating the association weights between each element in the sequence and other elements (a query-key-value mechanism). Furthermore, its multi-head mechanism splits the input vector into multiple subspaces (heads), each of which independently learns different attention patterns (e.g., the structural / semantic focus of time series x frequency domain features), and finally, the results are concatenated and linearly transformed. This module transforms the feature vector sequences of each node into a contextual representation containing spatiotemporal correlations through learnable spatiotemporal location encoding. Specifically, a cross-modal attention layer adaptively establishes semantic associations between acoustic event types and spatial spectral features, while a hierarchical attention mechanism effectively integrates local node features with global topological information. Finally, a fully connected network jointly optimizes the event localization coordinate estimation and type recognition tasks, employing a weighted loss function of mean squared error and cross-entropy for end-to-end training. This deep fusion framework provides more accurate positioning technology for intelligent sensing in the IoT environment.
[0073] This invention provides an acoustic event localization method applied to an acoustic event localization system, which includes a data processing subsystem and multiple high-order surround acoustic spherical array nodes. The method includes: the high-order surround acoustic spherical array nodes acquiring multimodal fusion data; wherein the multimodal fusion data includes: sound field information of the acoustic event, time information of the acoustic event, position information of the high-order surround acoustic spherical array nodes, attitude information of the high-order surround acoustic spherical array nodes, and inference type of the acoustic event; the high-order surround acoustic spherical array nodes compress the multimodal fusion data and then send it to the data processing subsystem; the data processing subsystem... The system extracts and identifies features from multimodal fusion data based on a pre-trained multi-source data fusion localization model, outputting the type and coordinates of acoustic events. The multi-source data fusion localization model includes a feature extraction module and a spatiotemporal fusion module. The feature extraction module includes temporal and spatial branches, while the spatiotemporal fusion module is a deep learning model based on a multi-head attention mechanism. By acquiring accurate multimodal fusion data and utilizing the localization model with both feature extraction and spatiotemporal fusion modules for feature extraction and identification, the system accurately outputs the type and location of acoustic events, simplifying computation and significantly improving the model's phenotypic performance even in complex environments.
[0074] Example 2
[0075] Based on the above embodiments, this invention provides an acoustic event localization device applied to an acoustic event localization system. The acoustic event localization system includes: a data processing subsystem and multiple high-order surround acoustic spherical array nodes; see also Figure 6 The diagram shown is a structural schematic of an acoustic event localization device provided in an embodiment of the present invention. The device includes:
[0076] The data acquisition module 310 is used to acquire multimodal fusion data from the high-order surround acoustic spherical array node; wherein, the multimodal fusion data includes: acoustic field information of acoustic events, time information of acoustic events, position information of the high-order surround acoustic spherical array node, attitude information of the high-order surround acoustic spherical array node, and inference type of acoustic events.
[0077] The data transmission module 320 is used to compress the multimodal fusion data from the high-order surround acoustic spherical array node and send it to the data processing subsystem.
[0078] The data analysis module 330 is used by the data processing subsystem to extract and identify features from multimodal fusion data based on a pre-trained multi-source data fusion localization model, and output the type and coordinates of acoustic events. The multi-source data fusion localization model includes a feature extraction module and a spatiotemporal fusion module. The feature extraction module includes a temporal branch and a spatial branch. The spatiotemporal fusion module is a deep learning model based on a multi-head attention mechanism.
[0079] Furthermore, in some preferred embodiments of the present invention, the high-order surround acoustic spherical array node includes: a spherical microphone array, a real-time dynamic differential system, an attitude sensor, and a sound field information processor; the data acquisition module 310 is used to acquire sound field information of acoustic events based on the spherical microphone array; acquire position information of the high-order surround acoustic spherical array node based on the real-time dynamic differential system; acquire attitude information of the high-order surround acoustic spherical array node based on the attitude sensor; and determine the inference type of the acoustic event based on the sound field information processor.
[0080] Furthermore, in some preferred embodiments of the present invention, the acoustic event localization system further includes: a meteorological monitoring node; the device further includes: a meteorological data acquisition module, used by the meteorological monitoring node to acquire meteorological information; wherein, the meteorological information includes at least one of the following: wind direction, wind speed, temperature, relative humidity, and air pressure; the meteorological monitoring node sends the meteorological information to the data processing subsystem.
[0081] Furthermore, in some preferred embodiments of the present invention, the device further includes: a meteorological data correction module, used by the data processing subsystem to correct the sound field information based on meteorological information.
[0082] Furthermore, in some preferred embodiments of the present invention, the high-order surround acoustic spherical array node transmits the compressed multimodal fusion data to the data processing subsystem through low-power wide-area transmission technology.
[0083] Furthermore, in some preferred embodiments of the present invention, the temporal branch captures the temporal correlation characteristics of the multimodal fusion data through a bidirectional gated logic unit; the spatial branch extracts the directional distribution features of the multimodal fusion data using a two-dimensional convolutional kernel.
[0084] Furthermore, in some preferred embodiments of the present invention, the spatiotemporal fusion module splits the input vector into multiple subspaces, then concatenates the results and performs a linear transformation to obtain the output result; wherein, the attention patterns of each subspace are different.
[0085] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the acoustic event localization device described above can be referred to the corresponding process in the embodiments of the aforementioned acoustic event localization method, and will not be repeated here.
[0086] Example 3
[0087] This invention also provides an acoustic event localization system for running an acoustic event localization method; see [link to related documentation]. Figure 7The diagram shown is a structural schematic of an acoustic event localization system provided by an embodiment of the present invention. The acoustic event localization system includes a memory 400 and a processor 401. The memory 400 is used to store one or more computer instructions, which are executed by the processor 401 to implement the acoustic event localization method described above.
[0088] Furthermore, Figure 7 The acoustic event localization system shown also includes a bus 402 and a communication interface 403. The processor 401, the communication interface 403, and the memory 400 are connected via the bus 402.
[0089] The memory 400 may include high-speed random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Communication between this system network element and at least one other network element is achieved through at least one communication interface 403 (which can be wired or wireless), such as the Internet, wide area network, local area network, or metropolitan area network. The bus 402 may be an ISA bus, PCI bus, or EISA bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 7 The symbol is represented by a single double-headed arrow, but this does not mean that there is only one bus or one type of bus.
[0090] Processor 401 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of processor 401 or by instructions in software form. Processor 401 can be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it can also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this invention. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this invention can be directly manifested as execution by a hardware decoding processor, or execution by a combination of hardware and software modules in the decoding processor. The software module can reside in a readily available storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory 400, and processor 401 reads information from memory 400 and, in conjunction with its hardware, completes the steps of the method described in the foregoing embodiments.
[0091] This invention also provides a computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are invoked and executed by a processor, they cause the processor to implement the aforementioned acoustic event localization method. For specific implementation details, please refer to the method embodiments, which will not be repeated here.
[0092] The computer program products of the acoustic event localization method, apparatus and system provided in the embodiments of the present invention include a computer-readable storage medium storing program code. The instructions included in the program code can be used to execute the methods in the preceding method embodiments. For specific implementation, please refer to the method embodiments, which will not be repeated here.
[0093] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the system and / or device described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0094] Furthermore, in the description of the embodiments of the present invention, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art can understand the specific meaning of the above terms in the present invention based on the specific circumstances.
[0095] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, essentially, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0096] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for acoustic event localization, characterized in that, An acoustic event localization system is applied, the acoustic event localization system comprising: a data processing subsystem and multiple high-order surround acoustic spherical array nodes; the method comprises: The high-order surround acoustic spherical array node acquires multimodal fusion data; wherein, the multimodal fusion data includes: acoustic field information of the acoustic event, time information of the acoustic event, position information of the high-order surround acoustic spherical array node, attitude information of the high-order surround acoustic spherical array node, and inference type of the acoustic event; The high-order surround acoustic spherical array node compresses the multimodal fusion data and sends it to the data processing subsystem. The data processing subsystem performs feature extraction and identification on the multimodal fusion data based on a pre-trained multi-source data fusion localization model, and outputs the type and coordinates of the acoustic event; wherein, the multi-source data fusion localization model includes: a feature extraction module and a spatiotemporal fusion module; the feature extraction module includes: a temporal branch and a spatial branch; the spatiotemporal fusion module is a deep learning model based on a multi-head attention mechanism; The high-order surround acoustic spherical array node includes: a spherical microphone array, a real-time dynamic differential system, an attitude sensor, and a sound field information processor; the steps for the high-order surround acoustic spherical array node to acquire multimodal fusion data include: Acquire the sound field information of the acoustic event based on the spherical microphone array; The position information of the high-order surround acoustic spherical array nodes is obtained based on the real-time dynamic differential system. The attitude information of the high-order surround acoustic spherical array nodes is obtained based on the attitude sensor; The inference type of the acoustic event is determined based on the sound field information processor.
2. The acoustic event localization method according to claim 1, characterized in that, The acoustic event localization system further includes: a meteorological monitoring node; the method further includes: The meteorological monitoring node acquires meteorological information; wherein the meteorological information includes at least one of the following: wind direction, wind speed, temperature, relative humidity, and air pressure; The meteorological monitoring node sends the meteorological information to the data processing subsystem.
3. The acoustic event localization method according to claim 2, characterized in that, The method further includes: The data processing subsystem corrects the sound field information based on the meteorological information.
4. The acoustic event localization method according to claim 1, characterized in that, The high-order surround acoustic spherical array node transmits the compressed multimodal fusion data to the data processing subsystem via low-power wide-area transmission technology.
5. The acoustic event localization method according to claim 1, characterized in that, The temporal branch captures the temporal correlation characteristics of the multimodal fusion data through a bidirectional gated logic unit; the spatial branch extracts the directional distribution features of the multimodal fusion data using a two-dimensional convolutional kernel.
6. The acoustic event localization method according to claim 1, characterized in that, The spatiotemporal fusion module splits the input vector into multiple subspaces, then concatenates the results and performs a linear transformation to obtain the output result; among them, the attention patterns of each subspace are different.
7. An acoustic event localization device, characterized in that, An acoustic event localization system is applied to the system, which includes a data processing subsystem and multiple high-order surround acoustic spherical array nodes; the device includes: The data acquisition module is used by the high-order surround acoustic spherical array node to acquire multimodal fusion data; wherein, the multimodal fusion data includes: sound field information of acoustic events, time information of acoustic events, position information of the high-order surround acoustic spherical array node, attitude information of the high-order surround acoustic spherical array node, and inference type of acoustic events; The data transmission module is used to compress the multimodal fusion data by the high-order surround acoustic spherical array node and then send it to the data processing subsystem. The data analysis module is used by the data processing subsystem to extract and identify features from the multimodal fusion data based on a pre-trained multi-source data fusion localization model, and output the type and coordinates of the acoustic event. The multi-source data fusion localization model includes a feature extraction module and a spatiotemporal fusion module. The feature extraction module includes a temporal branch and a spatial branch. The spatiotemporal fusion module is a deep learning model based on a multi-head attention mechanism. The high-order surround acoustic spherical array node includes: a spherical microphone array, a real-time dynamic differential system, an attitude sensor, and a sound field information processor; a data acquisition module is used to acquire sound field information of the acoustic event based on the spherical microphone array; acquire position information of the high-order surround acoustic spherical array node based on the real-time dynamic differential system; acquire attitude information of the high-order surround acoustic spherical array node based on the attitude sensor; and determine the inference type of the acoustic event based on the sound field information processor.
8. An acoustic event localization system, characterized in that, The method includes a processor and a memory, the memory storing computer-executable instructions that can be executed by the processor, the processor executing the computer-executable instructions to implement the acoustic event localization method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when invoked and executed by a processor, cause the processor to implement the acoustic event localization method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Audiovisual event positioning method fusing self-supervised multi-modal features
CN115393968A