Far-field accidental sound source positioning method, system, medium and equipment

By combining multimodal environmental perception and fast TDOA estimation with a physically constrained neural network, the problems of wasted computational resources and poor environmental adaptability in far-field sporadic sound source localization are solved, achieving efficient and accurate localization results. It is particularly suitable for real-time localization of short-term sudden sound events such as firecrackers and car horns.

CN121633985APending Publication Date: 2026-03-10ARMY ENG UNIV OF PLA
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing sound source localization technologies suffer from severe waste of computational resources, insufficient real-time performance, poor environmental adaptability, low positioning accuracy, and lack of reliability assessment when dealing with sporadic far-field events. They also fail to effectively utilize multi-dimensional information for assisted positioning.

Method used

Multimodal environmental perception technology is adopted. Audio signals are collected by sensors and preprocessed. Combined with multi-band energy analysis and environmental type recognition, the detection threshold is dynamically adjusted. Audio and environmental features are fused, and localization is performed using fast TDOA estimation and physical constraint neural networks. Confidence assessment is combined to improve localization accuracy and reliability.

Benefits of technology

It achieves efficient and accurate localization of far-field sporadic sound sources in complex urban environments, reduces computational load, and improves positioning accuracy and reliability. It is suitable for real-time localization of short-term sudden sound events such as firecracker sounds and car horns.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121633985A_ABST
    Figure CN121633985A_ABST
Patent Text Reader

Abstract

The invention discloses a far-field accidental sound source positioning method, system, medium and equipment, and the method comprises the steps: collecting an audio signal through a sensor, and carrying out the preprocessing of the audio signal, and obtaining a preprocessing signal; performing multi-band energy analysis to calculate short-time energy of each band, updating background energy estimation, and dynamically adjusting a detection threshold value; meanwhile, the processed audio features and real-time environment features are fused into multi-modal features, the multi-modal features are input into an accidental event detector for detection statistic calculation, and when the detection statistic reaches a self-adaptive threshold value, an accidental event detection stage is entered: a time window containing an accidental event in the preprocessed signal is determined, extended interception is carried out, and a time window containing the accidental event is obtained; the signals of the sensors in the corresponding time periods are extracted; fast TDOA estimation is carried out, abnormal estimation values are eliminated, and an estimation result is obtained; calculating the three-dimensional coordinates of the sound source by combining the position of the sensor; and calculating positioning confidence evaluation to obtain an effective positioning result. According to the invention, adaptive sound source localization of a complex urban environment is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to a method, system, medium, and device for locating far-field sporadic sound sources, belonging to the field of acoustic signal processing technology. Background Technology

[0002] Sound source localization technology has significant application value in fields such as urban noise monitoring, security early warning, and smart city construction. Traditional sound source localization methods mainly include geometric localization methods based on time difference of arrival (TDOA) and beamforming methods.

[0003] The existing technology has the following shortcomings:

[0004] 1. Severe waste of computing resources: Traditional methods adopt a continuous monitoring mode, resulting in more than 95% of invalid computation for sporadic events that account for only 0.1-5% of the total time; Insufficient short-time signal processing capability: Existing methods are mostly based on long-time average features, which have poor processing effect on sporadic events with a duration of 50ms-3s, resulting in low positioning accuracy.

[0005] 2. Poor environmental adaptability: In complex urban environments, background noise, multipath reflections and other interference factors severely affect positioning performance;

[0006] 3. Insufficient real-time performance: Traditional methods have high computational complexity and cannot meet the requirement of rapid location of sporadic events within 200ms;

[0007] 4. Lack of reliability assessment: Existing systems cannot provide confidence assessment of positioning results, which affects the reliability of practical applications.

[0008] 5. Lack of environmental awareness: Traditional methods cannot adaptively adjust detection strategies according to different environmental types (urban, industrial, natural, etc.), resulting in performance degradation when the environment changes;

[0009] 6. Limitations of single-modal information: It relies solely on audio signals and does not fully utilize multi-dimensional information such as environmental acoustic features and meteorological conditions for auxiliary positioning.

[0010] Currently, there are many patents related to sound source localization, most of which focus on near-field continuous sound localization technology. Some patents combine image and sound information for sound localization, but few address far-field sporadic event localization. Patent CN202410086072.6, "A Sound Source Localization and Event Detection Method Based on Deep Learning in Noisy Environments," proposes a deep learning-based method for sound source localization and event detection. However, its core focus is on processing noisy signals in high-noise environments, emphasizing the enhancement of sound features. This approach increases computational complexity and has limitations in responding to sporadic sound sources, making it unsuitable for real-time sound detection. Furthermore, the method's handling of overlapping sound sources is not ideal, limiting localization accuracy. Patent CN202510308744.8, "A ResNet Sound Source Localization Method Based on Attention Mechanism," proposes an improved ResNet sound source localization method that extracts phase differences in the frequency domain for sound source localization. However, it performs well with continuous sound sources, but its system complexity is high and it lacks flexibility. Patent CN202510715166.X: Sound Source Direction Finding and Localization Method and System for Complex Environments. The aforementioned patent proposes a sound source direction finding and localization method for complex environments, which relies on the audio similarity and sound intensity distribution between microphones to infer the direction of the sound source. However, this method has shortcomings in terms of computational complexity and practicality, with slow processing speed and difficulty in dealing with the high dynamic changes of short-term sporadic sound sources.

[0011] Therefore, there is an urgent need for an efficient and accurate localization method specifically designed for the characteristics of sporadic sound sources. Summary of the Invention

[0012] The purpose of this invention is to provide a method, system, medium, and device for locating far-field sporadic sound sources. Through multimodal environmental perception technology, it achieves adaptive sound source localization in complex urban environments, and is particularly suitable for rapid and accurate localization of short-term sudden sound events such as firecrackers and car horns.

[0013] To achieve the above objectives, the present invention is implemented using the following technical solution.

[0014] In a first aspect, the present invention provides a method for locating far-field sporadic sound sources, comprising:

[0015] Audio signals are collected by sensors, preprocessed, and a preprocessed signal is obtained.

[0016] Based on the preprocessed signal, multi-band energy analysis is performed to calculate the short-time energy of each frequency band. The background energy is updated by combining the environmental type and background noise, and the detection threshold is dynamically adjusted. Simultaneously, the processed audio features are fused with real-time environmental features to form multimodal features, and these multimodal features are input into the incidental event detector for detection statistics calculation. When the detection statistics reach an adaptive threshold, the incidental event detection stage begins.

[0017] The time window containing sporadic events in the preprocessed signal is determined, expanded and truncated, and the signals of each sensor corresponding to the time period are extracted.

[0018] Based on the signal, perform fast TDOA estimation, remove outlier estimates, and obtain the estimation result;

[0019] The three-dimensional coordinates of the sound source are calculated based on the estimation results and the sensor location;

[0020] Based on the three-dimensional coordinates, the location reliability assessment is calculated to obtain an effective location result.

[0021] Furthermore, the audio signal is synchronously acquired through four or more sensors, and the signals of each channel are pre-emphasized and normalized. A rectangular array configuration is used, and the sampling frequency is 16kHz.

[0022] The preprocessed signal undergoes multi-band energy analysis, decomposing it into low-frequency, mid-frequency, and high-frequency bands; the low-frequency band has a frequency range of 50-500Hz; the mid-frequency band has a frequency range of 500-2000Hz; and the high-frequency band has a frequency range of 2000-8000Hz.

[0023] Furthermore, based on long-term acoustic statistical characteristics, environmental types are identified and categorized into densely populated urban areas, industrial areas, and natural environments. Environmental parameters affecting sound propagation, such as humidity, temperature, and wind speed, are obtained in real time based on meteorological data.

[0024] Furthermore, the detection threshold is dynamically adjusted based on the environmental type and background noise, and a recursive strategy is used for updating, as expressed in the following expression:

[0025] ;

[0026] in, Forgetting factor, This is the minimum energy value for the current frame.

[0027] Furthermore, the short-time energy of each frequency band is calculated, and the expression is as follows:

[0028] ;

[0029] in, For frequency band, The length of the window.

[0030] Furthermore, update the background energy estimate and calculate the detection statistic, expressed as:

[0031] ;

[0032] in, For frequency band weights, For the first Frequency band energy, The detection threshold;

[0033] When the detection statistic reaches the adaptive threshold Then it enters the occasional event detection phase.

[0034] In the method of this invention, the location process is only entered when an occasional event is detected through an event-driven architecture, which significantly reduces the consumption of computational resources compared to the continuous location method that runs continuously.

[0035] Furthermore, audio features and environmental features are fused using an attention mechanism to obtain multimodal features, expressed as follows:

[0036] ;

[0037] in, For audio feature queries, As an environmental feature, Scaling factor Environmental characteristics The scaling operation is used to normalize the attention score, preventing excessively large values ​​from causing gradient vanishing and ensuring model training stability.

[0038] Furthermore, the start and end times of the time window containing the incidental event in the audio signal are determined using... This indicates that extended truncation is performed using... It indicates and extracts the signals of each sensor for the corresponding time period.

[0039] Furthermore, based on the signal, a fast TDOA estimation is performed. The fast TDOA estimator adopts a Physical Information Neural Network (PINN) architecture, with a one-dimensional convolutional neural network as the backbone. The input layer receives dual-channel audio signals, which are processed through three layers of convolution and an attention mechanism is embedded to separate multi-source signals. Physical sound propagation models (such as sound speed constraints and geometric equations) are incorporated into the training process. Finally, the TDOA estimate and separation vector are output through a fully connected layer.

[0040] The neural network architecture designed for the characteristics of sporadic signals in this invention significantly improves the accuracy of localization error compared to traditional methods under conditions of high signal-to-noise ratio.

[0041] Furthermore, based on the estimation results and the three-dimensional coordinates of the sound source calculated by the sensor, the following steps are taken: express.

[0042] Furthermore, the three-dimensional coordinates of the sound source are calculated by a position solver. The position solver, based on the spatial geometry of the sensor array, solves for the sound source position using the least squares method. The constraint expression is as follows:

[0043]

[0044] in, For geometric constraint loss function, For sensors and The measured TDOA between them Let be the three-dimensional position coordinates of the sound source to be solved. and The first One and The three-dimensional position coordinates of each sensor From sound source to sensor Euclidean distance, The speed of sound.

[0045] Furthermore, during the training phase, the physical laws of sound wave propagation are incorporated into the loss function and co-optimized with the PINN architecture of the fast TDOA estimator. The expression for the physically constrained loss function is as follows:

[0046] ;

[0047] in, The loss is data-driven (the difference between the true label and the predicted value). The physical constraint loss function is expressed as follows: , As a weighting factor, For the speed of sound, Sensors predicted by PINN network and Time Delay Amount (TDOA) between them and The first One and The three-dimensional position coordinates of each sensor The actual sound source locations in the training samples. From real sound source to sensor The Euclidean distance is used to ensure that the time delay difference predicted by the network conforms to the physical laws of sound wave propagation.

[0048] Furthermore, the location reliability assessment is calculated using the following expression:

[0049] ;

[0050] Where C represents the location reliability score, and Confidence_Net is the confidence evaluation network. This is the estimated value of TDOA. Here are the estimated coordinates of the sound source location, and SNR is the signal-to-noise ratio.

[0051] Furthermore, the confidence assessment comprehensively considers signal-to-noise ratio, TDOA variance, geometric factor, and spectral characteristic indicators.

[0052] when At that time, output a valid positioning result.

[0053] The method of this invention provides a confidence assessment mechanism, which effectively identifies and eliminates unreliable positioning results, improves the overall credibility of the system, and has a small total delay from event detection to location output, meeting the requirements of real-time applications.

[0054] In a second aspect, the present invention provides a far-field sporadic sound source localization system, comprising:

[0055] The data acquisition module is used to collect audio signals through sensors, perform preprocessing, and obtain preprocessed signals;

[0056] The incidental event detection module is used to perform multi-band energy analysis to calculate the short-time energy of each frequency band based on the preprocessed signal, update the background energy by combining environmental type and background noise, and dynamically adjust the detection threshold. Simultaneously, it fuses the processed audio features with real-time environmental features into multimodal features, and inputs these multimodal features into the incidental event detector for detection statistics calculation. When the detection statistics reach an adaptive threshold, the incidental event detection stage begins.

[0057] The signal interception module is used to determine the time window in the preprocessed signal that contains occasional events, perform extended interception, and extract the signals of each sensor for the corresponding time period;

[0058] The TDOA estimation module is used to perform fast TDOA estimation based on the signal, remove outlier estimates, and obtain the estimation result.

[0059] The position calculation module is used to calculate the three-dimensional coordinates of the sound source based on the estimation results and the sensor position;

[0060] The positioning evaluation module is used to calculate the positioning reliability evaluation based on the three-dimensional coordinates to obtain a valid positioning result.

[0061] Thirdly, the present invention provides a computer-readable storage medium having a computer program / instructions stored thereon, which, when executed by a processor, implements the steps of the far-field sporadic sound source localization method described in any of the first aspects.

[0062] Fourthly, the present invention provides a computer device, comprising:

[0063] Memory, used to store computer programs / instructions;

[0064] A processor for executing the computer program / instructions to implement the steps of the far-field sporadic sound source localization method as described in any one of the first aspects.

[0065] Compared with the prior art, the beneficial effects achieved by the present invention are as follows:

[0066] 1. The far-field sporadic sound source localization method provided by this invention collects audio signals from multiple sensors, performs normalization processing, and employs a sporadic event detector based on multi-band energy analysis and an adaptive threshold mechanism. It obtains multimodal features by fusing audio features and environmental features, then calculates detection statistics and checks whether an adaptive threshold is reached. For detection statistics that reach the threshold, the method proceeds to the sporadic event detection stage. An event window is determined and signal segments are extracted. A fast TDOA estimator is used for consistency verification, eliminating outlier estimates. In this invention, the fast TDOA estimator uses PINN, which effectively optimizes the time delay estimation of short-term sporadic signals by combining physical constraints and attention mechanisms, and supports multi-source separation. A sensor-to-position solver is used to calculate the three-dimensional coordinates of the sound source. The position solver is trained using collaborative constraints, further improving localization accuracy. Furthermore, a confidence estimator is used for localization, providing a reliability assessment of the localization results. Compared to traditional continuous localization methods, this invention reduces computation by more than 85%, making it particularly suitable for real-time localization applications of far-field sporadic sound sources such as firecrackers and car horns.

[0067] 2. The computer-readable storage medium and computer device provided by the present invention can execute the steps of the far-field sporadic sound source localization method provided by the present invention. Attached Figure Description

[0068] Figure 1 This is an overall structural diagram of the far-field sporadic sound source localization method provided according to an embodiment of the present invention;

[0069] Figure 2 This is a diagram showing the internal structure of the incidental event detector in the far-field incidental sound source localization method provided according to an embodiment of the present invention.

[0070] Figure 3This is a network architecture diagram of a fast TDOA estimator for a far-field sporadic sound source localization method provided according to an embodiment of the present invention;

[0071] Figure 4 This is a flowchart of a far-field sporadic sound source localization method provided in an embodiment of the present invention;

[0072] Figure 5 Comparison of time-spectral characteristics of sporadic events in different environmental types for the far-field sporadic sound source localization method provided in the embodiment of the present invention;

[0073] Figure 6 This is a multimodal structure diagram of the far-field sporadic sound source localization method provided according to an embodiment of the present invention. Detailed Implementation

[0074] It should be noted that:

[0075] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present invention and the specific features in the embodiments are detailed descriptions of the technical solution of the present invention, rather than limitations thereof. In the absence of conflict, the embodiments of the present invention and the technical features in the embodiments can be combined with each other.

[0076] The term "and / or" simply describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship.

[0077] Example 1

[0078] like Figure 1 As shown in the figure, this embodiment introduces a method for locating far-field sporadic sound sources, including:

[0079] Audio signals are collected by sensors, preprocessed, and a preprocessed signal is obtained.

[0080] Based on the preprocessed signal, multi-band energy analysis is performed to calculate the short-time energy of each frequency band. The background energy is updated by combining the environmental type and background noise, and the detection threshold is dynamically adjusted. Simultaneously, the processed audio features are fused with real-time environmental features to form multimodal features, and these multimodal features are input into the incidental event detector for detection statistics calculation. When the detection statistics reach an adaptive threshold, the incidental event detection stage begins.

[0081] The time window containing sporadic events in the preprocessed signal is determined, expanded and truncated, and the signals of each sensor corresponding to the time period are extracted.

[0082] Based on the signal, perform fast TDOA estimation, remove outlier estimates, and obtain the estimation result;

[0083] The three-dimensional coordinates of the sound source are calculated based on the estimation results and the sensor location;

[0084] Based on the three-dimensional coordinates, the location reliability assessment is calculated to obtain an effective location result.

[0085] This embodiment uses an embedded computing platform at the hardware level, equipped with an ARM Cortex-A78 processor and a Mali-G78 GPU, and a real-time audio processing framework at the software level, optimized based on PyTorch Mobile.

[0086] like Figure 2 As shown, in this embodiment, the audio signal is acquired through four omnidirectional condenser microphone sensors with a frequency response range of 20Hz-20kHz and a sensitivity of -38dB. A multi-channel audio acquisition card is used, and the synchronous sampling accuracy is [not specified]. The sensors are arranged in a rectangular array with the following coordinates:

[0087] Sensor 1: Coordinates (0, 0, 1.5m);

[0088] Sensor 2: Coordinates (5, 0, 1.5m);

[0089] Sensor 3: Coordinates (5, 5, 1.5m);

[0090] Sensor 4: Coordinates (0, 5, 1.5m);

[0091] In this embodiment, the sampling frequency is 48kHz and the bit depth is 16bit;

[0092] The preprocessed audio signal undergoes multi-band energy analysis. In this embodiment, a learnable digital filter bank is used to decompose the audio into low-frequency, mid-frequency, and high-frequency bands. The low-frequency band has a frequency range of 50-500Hz; the mid-frequency band has a frequency range of 500-2000Hz; and the high-frequency band has a frequency range of 2000-8000Hz. The bandpass filter is set to 50-8kHz, AGC is used for gain control, and noise reduction processing is performed.

[0093] Furthermore, such as Figure 5 As shown, the time-spectrum characteristics of different types of sporadic events were compared. In this embodiment, the environmental type was identified based on long-term acoustic statistical characteristics. The environmental type was divided into densely populated urban areas, industrial areas and natural environments. Environmental parameters such as humidity, temperature and wind speed that affect sound propagation were obtained in real time based on meteorological data, and a long time window (5-10 minutes) was used to analyze the environmental acoustic characteristics.

[0094] Furthermore, a sliding window method is used to calculate the short-time energy of each frequency band. The window size is 64ms (1024 samples @ 16kHz), the frame shift is 16ms (256 samples), and the window function used is the Hanning window, with the following expression:

[0095] ;

[0096] in, For frequency band, For frame shift, For the length of the window, For window functions.

[0097] Furthermore, the minimum tracking algorithm is used to estimate the background noise in each frequency band, and the background energy estimate is updated accordingly.

[0098] Furthermore, the detection threshold is dynamically adjusted based on the environmental type and background noise, and a recursive strategy is used for updating, as expressed in the following expression:

[0099] ;

[0100] in, Forgetting factor, This is the minimum energy value for the current frame.

[0101] Furthermore, such as Figure 6 As shown, audio features and environmental features are fused to obtain multimodal features, which are then input into an incidental event detector for statistical analysis. The incidental event detector uses learnable frequency band weights and obtains the optimal weight allocation for different event types through neural network training. In this embodiment, the calculation expression for the learnable weights is:

[0102] ;

[0103] in, These are trainable parameters, updated via backpropagation.

[0104] Furthermore, the detection statistic is calculated, and the final expression for the detection statistic is:

[0105] ;

[0106] in, Detection thresholds for each frequency band;

[0107] When the detection statistics reach the adaptive threshold, the event detection phase begins.

[0108] Furthermore, the start and end times of the time window containing the incidental event in the audio signal are determined using... In this embodiment, the window length is set to 1024, and the window is expanded and truncated with a repetition rate of 50%. It indicates and extracts the signals of each sensor for the corresponding time period.

[0109] Furthermore, such as Figure 3 As shown, based on the signal, fast TDOA estimation is performed. The fast TDOA estimator adopts the PINN architecture. In this embodiment, a one-dimensional convolutional neural network is used as the backbone. The input layer receives dual-channel audio signals, which are processed through three convolutional layers (the number of channels is set to 2→16→32→16). An attention mechanism is embedded to separate multi-source signals and to separate the physical sound propagation model (such as...). ,in For the speed of sound, For rapid TDOA, For sensor position, The sound source location is incorporated into the training process, and finally, the TDOA estimate and separation vector are output through a fully connected layer; the fast TDOA training loss function expression is:

[0110] ;

[0111] in, This is the physically possible maximum TDOA value.

[0112] Furthermore, such as Figure 4 As shown, the three-dimensional coordinates of the sound source are calculated based on the estimation results and the sensor position. express.

[0113] Furthermore, the three-dimensional coordinates of the sound source are calculated using a position coordinate analyzer. The position solver, based on the spatial geometry of the sensor array, solves for the sound source position using the least squares method. The constraint expression is as follows:

[0114]

[0115] in, For geometric constraint loss function, For sensors and The measured time delay difference (TDOA) between them Let be the three-dimensional position coordinates of the sound source to be solved. and The first One and The three-dimensional position coordinates of each sensor From sound source to sensor Euclidean distance, The speed of sound.

[0116] Furthermore, during the training phase, the physical laws of sound wave propagation are incorporated into the loss function and co-optimized with the PINN architecture of the fast TDOA estimator. The expression for the physically constrained loss function is as follows:

[0117] ;

[0118] in, For data-driven loss, The physical constraint loss function is expressed as follows: , As a weighting factor, For the speed of sound, Sensors predicted by PINN network and Time Delay Amount (TDOA) between them and The first One and The three-dimensional position coordinates of each sensor The actual sound source locations in the training samples. From real sound source to sensor The Euclidean distance.

[0119] Furthermore, the location reliability assessment is calculated using the following expression:

[0120] ;

[0121] Where C represents the location reliability score, and Confidence_Net is the confidence evaluation network. This is the estimated value of TDOA. Here are the estimated coordinates of the sound source location, and SNR is the signal-to-noise ratio.

[0122] The confidence assessment comprehensively considers signal-to-noise ratio (SNR), TDOA variance, geometric factor, and spectral characteristic indicators. In this embodiment, the SNR expression is: The expression for the variance of TDOA estimation is: The geometric factor expression is: The expression for spectral concentration is: .

[0123] when At that time, output a valid positioning result.

[0124] Example 2

[0125] Based on the far-field sporadic sound source localization method described in Embodiment 1, this embodiment introduces a far-field sporadic sound source localization system, including:

[0126] The data acquisition module is used to collect audio signals through sensors, perform preprocessing, and obtain preprocessed signals;

[0127] The incidental event detection module is used to perform multi-band energy analysis to calculate the short-time energy of each frequency band based on the preprocessed signal, update the background energy by combining environmental type and background noise, and dynamically adjust the detection threshold. Simultaneously, it fuses the processed audio features with real-time environmental features into multimodal features, and inputs these multimodal features into the incidental event detector for detection statistics calculation. When the detection statistics reach an adaptive threshold, the incidental event detection stage begins.

[0128] The signal interception module is used to determine the time window in the preprocessed signal that contains occasional events, perform extended interception, and extract the signals of each sensor for the corresponding time period;

[0129] The TDOA estimation module is used to perform fast TDOA estimation based on the signal, remove outlier estimates, and obtain the estimation result.

[0130] The position calculation module is used to calculate the three-dimensional coordinates of the sound source based on the estimation results and the sensor position;

[0131] The positioning evaluation module is used to calculate the positioning reliability evaluation based on the three-dimensional coordinates to obtain a valid positioning result.

[0132] Example 3

[0133] Based on the far-field sporadic sound source localization method described in Embodiment 1, this embodiment introduces a computer-readable storage medium storing a computer program / instruction thereon. When the computer program / instruction is executed by a processor, it implements the steps of the far-field sporadic sound source localization method as described in any of Embodiment 1.

[0134] Example 4

[0135] Based on the far-field sporadic sound source localization method described in Embodiment 1, this embodiment provides a computer device, including:

[0136] Memory, used to store computer programs / instructions;

[0137] A processor is configured to execute the computer program / instructions to implement the steps of the far-field sporadic sound source localization method as described in any one of Embodiment 1.

[0138] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0139] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0140] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0141] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0142] The embodiments of the present invention have been described above with reference to the accompanying drawings. However, the present invention is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of the present invention without departing from the spirit and scope of the claims. All of these forms are within the protection scope of the present invention.

Claims

1. A method of locating a far-field sporadic sound source, the method comprising: The method comprises the following steps: Collecting audio signals through sensors, preprocessing, and obtaining preprocessed signals; According to the preprocessed signals, the short-term energy of each frequency band is calculated by multi-band energy analysis, the background energy is updated combined with the environment type and background noise, and the detection threshold is dynamically adjusted; At the same time, the processed audio features and real-time environment features are fused into multi-modal features, and the multi-modal features are input into the incidental event detector for detection statistical quantity calculation, and when the detection statistical quantity reaches the adaptive threshold, it enters the incidental event detection stage: Determine the time window of the preprocessed signal containing the incidental event, perform extended interception, and extract the signal of each sensor corresponding to the time period; According to the signal, fast TDOA estimation is performed, and abnormal estimation values are removed to obtain the estimation result; According to the estimation result and the position of the sensor, the three-dimensional coordinates of the sound source are calculated; According to the three-dimensional coordinates, the position confidence evaluation is calculated to obtain the effective positioning result.

2. The far-field impulsive source localization method of claim 1, wherein, The audio signal is collected by 4 or more sensors, arranged in a rectangular array, and the sampling frequency is 16kHz; The preprocessed signal is analyzed by multi-band energy analysis and decomposed into low, medium and high frequency bands; The frequency range of the low frequency band is 50-500Hz; The frequency range of the medium frequency band is 500-2000Hz; The frequency range of the high frequency band is 2000-8000Hz.

3. The far-field impulsive source localization method of claim 1, wherein, The dynamic adjustment of the detection threshold adopts a recursive strategy for updating, and the expression is: ; wherein, is a forgetting factor, is the minimum energy value of the current frame.

4. The far-field impulsive source localization method of claim 1, wherein, The statistical detection quantity calculation formula is: ; wherein, is a frequency band weight, is a first is a frequency band energy, is a detection threshold.

5. The far-field impulsive source localization method of claim 1, wherein, The audio features and environment features are fused into multi-modal features using the attention mechanism, and the expression is: ; wherein, is an audio feature query, is an environmental feature, is a scaling factor, is an environmental feature is a dimension; The fast TDOA estimation uses a physical information neural network, the network input is a double-channel audio signal, the multi-source signal is separated by embedding the attention mechanism, and the output is the time delay difference estimation value and the separation result.

6. The far-field impulsive source localization method of claim 1, wherein, The three-dimensional coordinates of the sound source are calculated by a position solver, which is based on the spatial geometric relationship of the sensor array and solves the sound source position by the least square method, and the constraint condition expression is: ; in, The geometric constraint loss function, For sensors and The measured TDOA between them Let be the three-dimensional position coordinates of the sound source to be solved. and The first Individual and The three-dimensional position coordinates of each sensor From sound source to sensor Euclidean distance, The speed of sound.

7. The far-field impulsive source localization method of claim 1, wherein, The confidence evaluation includes signal-to-noise ratio, TDOA variance, geometric factor and spectral features.

8. A far-field sporadic sound source positioning system, characterized by, The method comprises the following steps: The data acquisition module is used to collect audio signals through sensors, preprocess, and obtain preprocessed signals; The incidental event detection module is used to calculate the short-term energy of each frequency band by multi-band energy analysis according to the preprocessed signals, update the background energy combined with the environment type and background noise, and dynamically adjust the detection threshold; At the same time, the processed audio features and real-time environment features are fused into multi-modal features, and the multi-modal features are input into the incidental event detector for detection statistical quantity calculation, and when the detection statistical quantity reaches the adaptive threshold, it enters the incidental event detection stage: The signal interception module is used to determine the time window of the preprocessed signal containing the incidental event, perform extended interception, and extract the signal of each sensor corresponding to the time period; The TDOA estimation module is used to perform fast TDOA estimation according to the signal, and remove abnormal estimation values to obtain the estimation result; The position solving module is used to calculate the three-dimensional coordinates of the sound source according to the estimation result and the position of the sensor; A positioning evaluation module is configured to calculate a positioning confidence evaluation according to the three-dimensional coordinates to obtain an effective positioning result.

9. A computer readable storage medium having stored thereon computer programs / instructions, characterized in that, The computer program / instructions, when executed by the processor, implement the steps of the far-field sporadic sound source positioning method in any one of claims 1 to 7.

10. A computer apparatus / device / system, characterized by, The computer program / instructions, when executed by the processor, implement the steps of the far-field sporadic sound source positioning method in any one of claims 1 to 7. The computer program / instructions, when executed by the processor, implement the steps of the far-field sporadic sound source positioning method in any one of claims 1 to 7. The computer program / instructions, when executed by the processor, implement the steps of the far-field sporadic sound source positioning method in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Sound source localization and event detection method in high-noise environment based on deep learning

    CN117953913A

  • ResNet sound source localization method based on attention mechanism

    CN120161409A

  • Sound source direction finding positioning method and system oriented to complex environment

    CN120233305A