Multi-source sound field positioning and separation algorithm based on deep neural network
By identifying source waveform segments in the audio stream and performing sliding cross-correlation operations, combined with fundamental frequency tracking and disorder evaluation, the problem of dynamic separation of direct sound and reflected sound in complex sound field environments is solved, and efficient sound source localization and separation is achieved, which is suitable for mobile devices.
Patent Information
- Application Number
- CN202511106087.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2045-08-08
AI Technical Summary
Existing technologies have difficulty achieving dynamic separation of direct sound and reflected sound in complex sound field environments, resulting in insufficient robustness in sound source localization. In addition, the reliance on high computing power and preset models limits real-time deployment on mobile devices.
By obtaining the energy changes of the audio stream, using the original waveform fragments as matching templates to perform sliding cross-correlation operations, identifying direct sound and reverberation, and combining fundamental frequency tracking and disorder evaluation mechanisms, the parallel positioning and separation of multiple sound sources can be achieved.
Without the need for high computing power and preset models, highly robust sound source localization and separation in complex sound field environments is achieved, reducing the computing load and improving the adaptability to dynamic environments.
Smart Images

Figure CN120595237A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a multi-source sound field localization and separation algorithm based on a deep neural network, belonging to the field of speech or sound processing technology. Background Art
[0002] In the field of speech or sound processing technology, the positioning and separation of multi-source sound fields are key technologies for realizing scenarios such as industrial equipment anomaly monitoring and intelligent cockpit interaction. Current mainstream solutions generally rely on deep neural networks to build end-to-end acoustic mapping models, and learn the relationship between sound source characteristics and spatial positions through massive data training. Although such methods can capture some acoustic laws, they have two fundamental limitations: First, the sound field is simplified into a linear superposition signal, ignoring the nonlinear interference effect caused by sound waves reflecting at rigid boundaries such as equipment casings and walls, resulting in feature confusion when sudden sound sources (such as metal collisions) are coupled with steady-state noise (such as engine hum); second, increasing network depth to address insufficient environmental adaptability forces computing power requirements to far exceed the carrying capacity of edge devices, making it difficult to deploy in real time in mobile scenarios such as industrial inspection robots.
[0003] Taking the forklift collision detection scenario as an example, existing technologies require a preset sound field environment geometric model to suppress reverberation interference. However, when sound waves pass through the shelf partition area, the difference in acoustic impedance between metal and cloth materials causes the reflection path calculation deviation to exceed 30%, forcing the system to frequently switch to the low-precision direct sound separation mode. The industry has tried to compensate for Doppler frequency shift by introducing high-precision motion sensors, but the added hardware costs and real-time calibration burden have further aggravated the contradiction between system complexity and power consumption.
[0004] A deeper technical contradiction lies in the fact that the existing paradigm treats reverberation as noise that needs to be eliminated rather than a physical messenger containing spatial information. This cognitive bias traps the system in a passive cycle of data fitting the environment. To improve accuracy, network parameters are forced to be stacked, but because the universal laws of acoustic boundary effects are ignored, the problem of robustness of sound source localization in dynamic environments cannot be fundamentally solved. Specifically, the existing technology has the following core bottlenecks: 1. Reliance on preset CAD models cannot respond to scene geometry changes (such as mobile device displacement) in real time; 2. Physical variables such as material acoustic impedance differences and sound source motion state are separated from signal processing, resulting in insufficient model generalization; 3. Complex models are seriously mismatched with edge hardware capabilities, restricting the implementation of technology. Therefore, how to analyze and reconstruct environmental cognition mechanisms through the autocorrelation characteristics of sound waves and achieve dynamic decoupling of multi-source sound fields without the need for high computing power and preset models has become the technical problem to be solved by this invention. Summary of the Invention
[0005] The present invention provides a multi-source sound field localization and separation algorithm based on deep neural network, the main purpose of which is to solve the problem of dynamic separation of direct sound and reflected sound in complex sound field environment.
[0006] To achieve the above objectives, the present invention provides a multi-source sound field localization and separation algorithm based on a deep neural network, the algorithm comprising the following steps: Step a: obtaining an audio stream and continuously monitoring the energy of the audio stream; when the energy of the audio stream first exceeds a dynamic energy threshold dynamically adjusted based on the real-time average energy of the ambient background noise, intercepting an initial audio segment of a predetermined length from that moment and determining it as a source waveform segment, where the source waveform segment is the initial waveform characteristic of the sound source itself; Step b: Using the source waveform segment as a matching template, a sliding cross-correlation operation is performed on the audio stream following the source waveform segment to detect one or more distortion-related peaks that are delayed in the time domain and are highly correlated with the waveform features of the source waveform segment and have energy attenuation. The time delay and amplitude attenuation of the distortion-related peaks are information about the path and environment of sound propagation. Step c, based on the original waveform segment and the audio signal that follows it, determines that it is the direct sound, and identifies the audio signal segment corresponding to one or more distortion correlation peaks as the reverberation of the direct sound; based on the identification results of the direct sound and the reverberation in the time domain, performs sound source localization processing on the audio stream; when an energy mutation is detected in the audio stream whose energy value exceeds the dynamic energy threshold and the cross-correlation with the currently used original waveform segment is lower than the predetermined cross-correlation threshold, a new original waveform segment is generated, and the steps of detecting distortion correlation peaks, determining direct sound and reverberation, and performing sound source localization processing are performed in parallel on the new original waveform segment, thereby realizing parallel localization and separation of multiple sound sources.
[0007] Preferably, after the distortion correlation peak detection step and before the sound source localization processing step, it also includes: continuously tracking the fundamental frequency of the original waveform segment and the audio signal that follows it to obtain the trajectory of the fundamental frequency changing with time; judging whether the fundamental frequency trajectory shows a continuous monotonic change; if it is judged as no, continuing to execute the sound source localization processing step; if it is judged as yes, terminating the execution of the sliding cross-correlation operation step, and switching to a collaborative tracking method based on energy envelope and arrival time difference, wherein the continuous energy envelope of the sound source is continuously tracked, and the arrival time difference of the energy envelope is calculated using a dual-microphone array to obtain the direction of the moving sound source.
[0008] Preferably, after the distortion correlation peak detection step and before the sound source localization processing step, it also includes: forming a reverberation feature sequence based on the occurrence time position and / or amplitude of one or more distortion correlation peaks in the time domain; calculating the disorder index of the reverberation feature sequence, the disorder index is determined according to the variance of the time interval sequence and the monotonicity degree of the amplitude attenuation sequence in the reverberation feature sequence; when the disorder index is lower than the predetermined disorder threshold, executing the sound source localization processing step; when the disorder index is higher than the predetermined disorder threshold, terminating the subsequent processing steps, and identifying the current source waveform segment as background noise.
[0009] Preferably, when the disorder index is higher than a predetermined disorder threshold and is identified as background noise, the current source waveform segment is marked as a background noise template, and in subsequent audio stream monitoring, when the energy exceeds the dynamic energy threshold, the background noise template is preferentially used to match the energy mutation. If the match is successful, no subsequent processing is performed.
[0010] Preferably, the predetermined length of the initial audio segment is five milliseconds to ten milliseconds.
[0011] Preferably, the sliding cross-correlation operation is accelerated by fast Fourier transform.
[0012] Preferably, the sound source localization processing includes: when a dual-microphone array is used, according to the time delay difference between the first detected distortion correlation peak and the two microphones of the dual-microphone array , calculate the azimuth of the reflecting surface Azimuth Calculated using the following formula: ,in, is the speed of sound, is the distance between the two microphones.
[0013] Preferably, the method further includes performing sound source separation processing on the audio stream, wherein the sound source separation processing includes: applying a predetermined gain suppression factor to perform gain suppression on the audio signal segment identified as reverberation.
[0014] Preferably, the method further includes performing sound source separation processing on the audio stream, where the sound source separation processing includes: subtracting the energy of the reverberation from the mixed audio signal.
[0015] Preferably, the method is executed by a microcontroller, and the microcontroller is configured to perform an energy detection operation, an audio slicing operation, and a one-dimensional cross-correlation operation.
[0016] Compared with the prior art, the present invention has the following beneficial effects: 1. By using the initial waveform of the sound source as an autocorrelation template, the system acquires for the first time the intrinsic ability to parse spatial information from reverberation. When sound waves propagate in a complex environment, the temporal correlation between the direct sound and its distortion-related peaks enables the system to directly utilize the path delay information of the reflected sound to naturally distinguish the sound source itself from environmental reflectors. This cognitive leap, transforming reverberation from an interference signal to an environmental messenger, avoids the complex calculations of traditional solutions to combat reverberation, and shifts sound field understanding from passive decomposition to active perception.
[0017] 2. The combination of a native waveform capture mechanism triggered by dynamic energy thresholds and multi-event parallel processing logic enables the system to automatically distinguish between sound source events of different natures through waveform correlation in scenarios where sudden sound sources such as metal collisions coexist with steady-state noise such as machine roars. The system no longer requires preset models for specific sound source types. Instead, it spontaneously forms multi-source processing channels based on the intrinsic time-domain fingerprint characteristics of sound waves, which in principle resolves the performance conflicts of traditional models when heterogeneous sound sources are mixed. When the continuous drift of the fundamental frequency trajectory reveals the motion state of the sound source, the system dynamically switches to energy envelope tracking mode. This autonomous evolution of processing logic eliminates the Doppler effect of the moving sound source, which no longer causes positioning failure, but instead transforms it into a natural indicator of motion detection. By reusing the arrival time difference calculation of the two microphones, the system maintains the continuity of sound source direction tracking while avoiding complex motion compensation models, achieving a natural unification of static analysis and dynamic tracking.
[0018] 3. The disorder evaluation mechanism of the reverberation peak sequence enables the system to identify the acoustic background. When a disordered correlation peak caused by steady-state noise is detected, the system automatically terminates invalid calculations and marks the current waveform as a noise template. This mechanism makes lightweight judgments on information, making the spectral characteristics of the background noise an immune marker for the system's self-learning, significantly reducing the false trigger rate while maintaining sensitivity to weak valid signals. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 Schematic diagram of the sound wave propagation path and reflection characteristics of the present invention; Figure 2 This is a timing diagram of the background noise template recognition and disorder determination process of the present invention; Figure 3 This is a comparison chart of the calculated values of the disorder index under different sound source types of the present invention.
[0020] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0021] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.
[0022] The present invention discloses a multi-source sound field localization and separation algorithm based on a deep neural network. Its overall architecture and operation process are mainly composed of core stages such as parallel capture of source waveforms, collaborative identification of sound source states, and temporal decoupling of sound field information. These stages work together to transform reverberation signals in complex environments from interference sources in traditional cognition into physical messengers carrying spatial information, thereby achieving highly robust localization and separation of multiple static and dynamic sound sources without the need for preset environmental models and high-computing power hardware.
[0023] Specifically, the typical application scenario is an intelligent inspection robot or automatic forklift deployed in a complex indoor environment such as a factory or warehouse. The core processor of the system is usually an embedded microcontroller, which continuously performs real-time acquisition and analysis of the audio stream. The microcontroller first continuously monitors the environmental background noise, and dynamically generates and adjusts a dynamic energy threshold by calculating the real-time average energy within a time window. The threshold is set to be slightly higher than the steady-state noise level of the current environment. This ensures that the system has a sensitive perception of energy mutations, while effectively avoiding false triggering caused by meaningless fluctuations in background noise. When a transient acoustic event occurs in the scene, such as the impact sound of a metal tool falling from a height, the energy of the sound signal will instantaneously and significantly exceed the aforementioned dynamic energy threshold. At this time, the system is immediately triggered, and from the moment the energy mutation occurs, the system is intercepted. An initial audio clip of a predetermined length is taken; the length of the clip is set between five and ten milliseconds. The selection of this length is based on a key technical trade-off: the duration of five to ten milliseconds is sufficient to fully capture the unique initial waveform fingerprint of a sound source such as a metal impact sound, but short enough to ensure with a high probability that the intercepted clip contains only direct sound without any reflection, thereby ensuring the purity of the template at the source; the intercepted initial audio clip is then determined by the system as a source waveform clip, which serves as an acoustic fingerprint template representing the physical characteristics of the sound source itself. After successfully capturing the source waveform clip, the system immediately enters the sliding cross-correlation operation stage, in which the source waveform clip is used as a matching template to perform continuous sliding matching calculations on the real-time audio stream that follows it; to ensure that the operation can be efficiently executed on a resource-constrained microcontroller, the sliding cross-correlation operation is performed through a fast Fourier transform It accelerates the computationally intensive convolution operation in the time domain by converting it into a more efficient multiplication operation in the frequency domain, greatly reducing the computational load. The core goal of this operation is to detect one or more delayed distortion-related peaks in the time domain that are highly correlated with the waveform characteristics of the original waveform fragment but attenuated in energy. These distortion-related peaks are essentially reverberations formed after the direct sound is reflected through different physical paths. Their time delay relative to the direct sound directly reflects the difference in the propagation path length of the sound wave, while the attenuation of their amplitude reflects the acoustic absorption characteristics of the reflection interface and the energy loss caused by the propagation distance. Therefore, the sequence composed of these distortion-related peaks contains rich information about the geometric structure of the sound field environment.
[0024] Based on the above calculation results, the system will identify the original waveform segment as the matching template and its first appearance position in the audio stream as the direct sound, and identify the audio signal segments corresponding to all subsequently detected distortion-related peaks as the reverberation of the direct sound; by analyzing the arrival time delay difference between the direct sound and the first strongest distortion-related peak, which is usually the reflected sound generated by the nearest single reflecting surface, between the dual-microphone array deployed on the robot, the system can perform preliminary sound source localization processing; specifically, if the fixed distance between the two microphones of the dual-microphone array is , the speed of sound is The system measures the time delay difference between the first detected distortion correlation peak reaching the two microphones as , the system can calculate the azimuth angle of the physical reflective surface that produces the reflection relative to the robot's own coordinate system according to the following formula : By calculating the orientation of one or more such reflecting surfaces, the system can reversely infer the possible location of the sound source, achieving high-precision localization. In addition to this basic localization process, this technical solution also designs a parallel processing and state switching mechanism to cope with complex acoustic scenarios. When the system is processing a sound source event, such as a metal impact sound, if another energy mutation occurs in the audio stream with an energy value exceeding the dynamic energy threshold, such as a staff member shouting nearby, the system will first calculate the cross-correlation between this new energy mutation waveform and the currently used source waveform segment. Because the waveform characteristics of the shout and the metal impact sound are completely different, their cross-correlation will be significantly lower than a preset cross-correlation threshold. Based on this judgment, the system determines that this is a new, independent sound source event. The system then generates a new source waveform segment for the shout and opens a parallel processing channel for this purpose, independently executing the subsequent steps of detecting distortion correlation peaks, determining direct sound and reverberation, and performing sound source localization processing. This achieves the parallel localization and separation of multiple sound sources of different origins and properties.
[0025] Furthermore, in order to handle special cases of moving sound sources, such as the engine sound of a forklift that is continuously approaching or moving away, the system also introduces a judgment logic for the motion state of the sound source after the distortion correlation peak detection step; the system will continuously track the fundamental frequency of the source waveform segment and the audio signal that follows it to obtain the trajectory of the fundamental frequency changing with time. If the fundamental frequency trajectory shows a clear and continuous monotonic change, whether it is a monotonic increase or a monotonic decrease, it constitutes a significant physical evidence of the Doppler effect, indicating that the sound source is moving radially relative to the robot; once the sound source is confirmed to be in the In a moving state, the system will determine that the reverberation-based positioning method may lose accuracy due to the rapid change of the sound source position; therefore, the system will terminate the subsequent sliding cross-correlation calculation steps and dynamically switch to a collaborative tracking method based on energy envelope and arrival time difference; in this method, the system will continue to track the continuous energy envelope of the moving sound source, and use the dual-microphone array to calculate the arrival time difference of the energy envelope between the two microphones in real time, so that it can continuously and stably track the direction of the moving sound source, realizing intelligent and seamless switching between static positioning and dynamic tracking modes.
[0026] In addition, in order to enhance the system's ability to identify and exempt environmental background noise, the system also includes a disorder evaluation mechanism based on reverberation feature sequences; after detecting a series of distortion-related peaks, the system will form a reverberation feature sequence based on the time position and amplitude of these correlation peaks in the time domain. Subsequently, the system will calculate the disorder index of the sequence, which combines two dimensions: the variance of the time interval sequence and the degree of monotonicity of the amplitude attenuation sequence; a reverberation sequence composed of clear physical reflections will usually show strong physical regularity in its arrival time interval and energy attenuation, resulting in a lower disorder index. On the contrary, if a sound source is essentially diffuse and continuous background noise such as a large fan or air-conditioning system, its autocorrelation operation may produce a series of seemingly related but time-series and Pseudo-peaks with chaotic energy attenuation patterns cause the disorder index to be higher than a predetermined disorder threshold. When this happens, the system determines that the current source waveform segment does not originate from a clear acoustic event, but from background noise, and thus terminates the subsequent positioning and separation processing steps. Thereafter, in order to further improve system efficiency, the system will mark the source waveform segment identified as background noise as a background noise template. In subsequent audio stream monitoring, when a new energy mutation exceeds the threshold, the system will give priority to using the stored background noise template to match the mutation. If the match is successful, it will be directly determined to be known background noise and ignored, and no subsequent processing will be initiated, thereby achieving adaptive immunity to specific background noise, significantly reducing the system's false trigger rate and unnecessary computational overhead.
[0027] Finally, after completing the positioning of the sound source, this technical solution also provides the function of sound source separation processing to improve the clarity of the target sound source; one separation method is to apply a predetermined gain suppression factor to all audio signal segments identified as reverberation, that is, significantly reduce their volume without completely eliminating them; another more direct method is to estimate and confirm the total energy of the reverberation part, and then subtract this part of the reverberation energy from the mixed overall audio signal. Through these methods, the interference of environmental reflections on the main sound source can be effectively suppressed, and a purer direct sound signal can be extracted, ultimately improving the application efficiency of the entire system.
[0028] Example 1: The technical solution of the present invention operates in a continuously running fully automatic three-dimensional warehouse environment in the following manner: a high-speed conveying and sorting system is deployed in the warehouse, and an unmanned forklift is operating between metal shelves. When a cargo box slides off the conveyor belt and hits the concrete floor, the energy of the short impact sound wave generated by it immediately exceeds the dynamic energy threshold that is dynamically adjusted according to the real-time average energy of the environmental background noise. The system then intercepts an initial audio segment of a predetermined length of five to ten milliseconds from that moment on, and determines it as the original waveform segment of the impact event; the system immediately uses this original waveform segment as a matching template for the subsequent The audio stream performs a sliding cross-correlation operation based on fast Fourier transform acceleration, and detects a series of distortion-related peaks in the time domain caused by sound waves hitting the surrounding metal shelves and walls. These peaks are highly correlated with the original waveform characteristics, but their energy decays sequentially. By transforming the reverberation signal from interference to be eliminated into an environmental messenger containing spatial information, the system fundamentally redefines the technical paradigm of sound field perception without relying on any preset warehouse geometry model, and transforms the understanding of the environment from passive model fitting to active analysis based on the autocorrelation characteristics of sound waves. While processing the above-mentioned collision event, the engine roar of an accelerating unmanned forklift was heard. The system processes the sound as an independent acoustic event in parallel. Because the cross-correlation between its waveform characteristics and the source waveform segment of the aforementioned impact sound is lower than the preset threshold, the system generates a new source waveform segment for it. After continuously tracking the fundamental frequency of the new segment, the system recognizes that its fundamental frequency trajectory shows a continuous monotonic change. This change reveals that the sound source is in motion and produces a Doppler effect. At this time, a synergistic mechanism is triggered: the sound source motion state obtained by fundamental frequency tracking provides a decisive priori condition for subsequent processing. Based on this, the system terminates the sliding cross-correlation operation for the sound source and switches to the synergistic operation based on energy envelope and arrival time difference. Using the same tracking method, a dual-microphone array calculates the time delay difference between the energy envelope reaching the two microphones to determine the direction of the sound source. This autonomous switching between static analysis and dynamic tracking modes matches the processing logic with the physical state of the sound source, significantly reducing the computational load while ensuring the continuity and accuracy of tracking sound sources of varying nature. During this process, the ventilation system on the warehouse roof continuously emits a steady low-frequency hum, and its energy fluctuations may briefly trigger the energy threshold. The system also generates source waveform segments for this hum and initiates cross-correlation detection. However, the resulting distorted correlation peak sequence exhibits a high degree of randomness in both its temporal distribution and amplitude decay.The system calculates the disorder index of the reverberation characteristic sequence—determined by the variance of the time interval sequence and the monotonicity of the amplitude decay sequence—and finds that its value exceeds a predetermined threshold. This mechanism physically resolves the inherent contradiction between increasing sensitivity and suppressing false triggers in the industry. The system does not rely on a voiceprint database, but instead, based on the inherent laws of sound wave propagation, clearly identifies the waveform as background noise and marks it as a background noise template. In subsequent monitoring, any energy mutations matching this template are directly ignored, thus achieving effective immunity to steady-state environmental noise without compromising the ability to capture true transient events.
[0029] The overall architecture of the algorithm, by internalizing the laws of acoustic physics into an analysis of signal timing and intrinsic correlation, makes the geometric structure of the sound field environment and the motion state of the sound source no longer external variables that require preset models to passively adapt, but endogenous information that can be perceived in real time through the signal's own evolutionary characteristics. This design avoids the traditional technical path's reliance on massive labeled data, complex neural network models, and high-computing-power hardware, allowing highly robust multi-source sound field positioning and separation functions to be implemented on a microcontroller that only needs to perform basic operations such as energy detection, audio slicing, and one-dimensional cross-correlation calculations.
[0030] Example 2: To quantitatively verify the core capability of the present invention in distinguishing and locating transient impact sound sources from steady-state background noise in complex sound fields, this example aims to accurately test the algorithm's ability to effectively capture burst signals, spatially locate them using their reverberation information, and autonomously identify and suppress coexisting steady-state noise, without a preset environmental model. This demonstrates the technical effectiveness of the present invention in converting reverberation from an interference signal into an environmental messenger and achieving autonomous collaborative processing of heterogeneous sound sources. To this end, a reproducible physical test platform simulating an industrial environment was constructed. The platform was built in a test room measuring 5 meters by 5 meters by 3 meters. Metal plates and sound-absorbing cotton were arranged alternately on the walls to simulate the highly complex acoustic reflection environment of a factory workshop where shelves and open spaces coexist. The acoustic sensing core was a dual-microphone array with a spacing of 15 cm, placed in a corner of the room as a signal receiver. The sound source consisted of two parts: one deployed at the indoor coordinates 4.0 and 3.5. An industrial-grade loudspeaker continuously plays a recording of a steady-state forklift engine noise with an average sound pressure level of 75 decibels. An electrically controlled air cannon, capable of generating instantaneous, high-intensity shocks with a peak sound pressure level of 110 decibels, can be triggered from anywhere in the room to simulate sudden events such as falling cargo or equipment collisions. A key parameter in the experiment is the predetermined length of the source waveform segment, which must strike a technical balance between ensuring waveform uniqueness and avoiding early reverberation contamination. The selection of this parameter directly affects the accuracy of the subsequent sliding cross-correlation operation and is a prerequisite for determining whether the system can successfully decouple the direct sound signature from the signal. The core of this technical trade-off is that if the segment is too short, it may not contain sufficient waveform detail to distinguish it from other acoustic events, resulting in a lower signal-to-noise ratio for the correlation analysis. Conversely, if the segment is too long, it is very likely to include primary reflections from the reflective surface closest to the sound source, thereby contaminating the source waveform used as the matching template and losing its purity as a representation of the direct sound. Therefore, the decision rule for this length is modeled as follows: its duration should be significantly longer than the energy rise time of the target transient sound source to ensure that the core features are captured, and at the same time it must be less than the first reflection delay determined by the minimum reflection path in the room to ensure the purity of the template. In this test environment, the main energy pulse of the air cannon impact sound is completed within 2 milliseconds, and the first reflection sound delay caused by the sound source being at least 1 meter away from the nearest wall is approximately Based on this decision model, a non-restrictive sample length of 8 milliseconds is set. This value not only completely covers the core waveform of the impact sound, but also effectively avoids the interference of early reflections, laying a solid foundation for subsequent autocorrelation analysis.
[0031] After the experiment was started, the system first operated in an environment with only the steady-state noise of a forklift engine. The algorithm detected that the energy of the audio stream continued to stably exceed the energy threshold dynamically generated based on the background silence level. It then intercepted the first 8-millisecond audio clip as the initial source waveform clip. The system used this as a template to perform a sliding cross-correlation operation on the subsequent audio stream. The key phenomenon observed at this time was that the correlation peak sequence showed a high-density, low-attenuation, and disordered form. That is, a large number of correlation peaks appeared densely, the variance of their time intervals was extremely large, and the amplitude decay sequence did not have a clear monotonicity. The system quantified this based on the reverberation characteristic sequence disorder evaluation mechanism of the present invention, as shown in Table 1.
[0032] Table 1: Characteristics of different sound sources and system judgment results.
[0033] As shown in the data in the table above, the waveform segment triggered by steady-state noise has a disorder index of 0.91 for its reverberation characteristic sequence, which is significantly higher than the system preset disorder threshold of 0.8. The inherent mechanism of this data result is that the reverberation sound field formed by steady-state noise in a complex environment is a random process of continuous superposition and energy diffusion. Its reflection path appears as a disordered arrival in the time domain, lacking the clear temporal structure from direct sound to each reflected sound of the instantaneous impact sound source; therefore, the algorithm successfully identifies and marks the original waveform segment as a background noise template based on this high disorder index, and terminates the invalid calculation of its positioning, thereby avoiding the waste of computing power. This directly confirms the use of the reverberation disorder evaluation mechanism of the present invention to identify and exempt background noise. The ability of sound. Subsequently, with the engine noise playing continuously in the background, the air cannon at coordinates 2.5, 2.0 was triggered. The system detected a dramatic energy mutation with an energy value far exceeding the dynamic threshold and a cross-correlation with the stored background noise template below the preset threshold of 0.3. Based on the multi-event parallel processing logic of the present invention, the system immediately determined this to be a new, independent acoustic event and generated a new 8-millisecond native waveform segment for it. The subsequent sliding cross-correlation calculation showed a phenomenon completely different from the previous one: a clear, direct sound signal with the highest energy, followed by a series of distortion-related peaks with delayed and regularly decaying energy in the time domain. The system used a dual-microphone array to accurately record the arrival time of these key signals (see Table 2).
[0034] Table 2: Time difference of arrival for different signal types.
[0035] The system uses the first reflection peak as a signal to analyze spatial information and uses the time delay difference between its arrival at the dual microphone array to Microseconds, calculate the azimuth of the reflecting surface according to the formula , the calculation process is The calculation results are consistent with the air cannon and the nearest reflecting surface, which is located at This process verifies the core mechanism of the present invention: by using the sound source's own waveform as a template for autocorrelation matching, the complex reverberation signal is successfully decoupled into a distorted correlation peak sequence containing propagation path and delay information, thereby converting the reverberation that is regarded as interference in traditional methods into key physical information for highly robust sound source localization.
[0036] Example 3: This example combines Figures 1 to 3 , this paper describes the implementation of a multi-source sound field localization and separation algorithm based on deep neural network. Figure 1 As shown in the figure, the horizontal axis is time (ms) and the vertical axis is amplitude. By marking and classifying the waveforms of multiple key signals in the process of sound wave propagation, the amplitude attenuation and time delay characteristics of the sound wave signal under different propagation paths are reflected. The first group of waveforms on the far left is represented by a thick solid line and is marked as the source waveform. It has the largest amplitude and represents the direct sound signal initially emitted by the sound source, corresponding to the line description of the direct sound in the legend; followed by As a mark, a group of waveforms represented by dotted lines appear, which are the first reflection signals, marked as the first reflection, representing the signal response after the sound wave is reflected by an obstacle once; followed by As intervals, a dotted line waveform appears, marked as secondary reflection, indicating the propagation signal of the sound wave after two reflections; Reflected waveforms continue to appear at intervals, with their amplitudes gradually weakening, indicating the gradual attenuation of signal energy due to multiple reflections.
[0037] like Figure 2 As shown, first, the main control module initiates an instruction to the reverberation analysis module to request analysis of the time-varying correlation peak sequence. After completing the analysis, the reverberation analysis module returns the reverberation feature sequence. Then the main control module requests to calculate the disorder based on the feature sequence. This request is sent to the disorder evaluation module. The disorder evaluation module evaluates the disorder of the waveform by calculating the time interval variance and the amplitude decay monotonicity, and returns the disorder index value to the main control module. The main control module determines whether the index is higher than the predetermined threshold based on the return value. If the result satisfies [the index is higher than the threshold], the main control module determines the current waveform as background noise, and continues to mark and store it as a background noise template. After the operation is completed, the main control module further confirms that the template has been stored, and executes the termination of subsequent positioning and separation processing, terminating the subsequent sound source positioning and signal separation process of the background noise waveform.
[0038] like Figure 3As shown, the horizontal axis is the calculated value, and the vertical axis is the monotonicity of the amplitude attenuation sequence, the variance of the time interval sequence, and the calculated disorder index. The three parameters are used to evaluate the attenuation regularity, time consistency, and overall disorder of the sound wave characteristics, respectively. The figure compares two types of sound sources, namely air cannon impact and engine steady-state noise, which are marked with different shapes in the legend. The circle represents the air cannon impact and the square represents the engine steady-state noise. It can be seen from the data in the figure that the calculated values of the air cannon impact in terms of the monotonicity of the amplitude attenuation sequence, the variance of the time interval sequence, and the calculated disorder index are significantly lower than those of the engine steady-state noise, especially in the dimension of the calculated disorder index. The corresponding calculated value of the air cannon impact is about 0.1, while the corresponding value of the engine steady-state noise is close to 0.9.
[0039] Example 4: In this example, when the system is first deployed in a specific working environment, a single environmental self-calibration process is performed to set the predetermined length of the source waveform segment. After the process is started, the system drives a miniature acoustic exciter configured near its dual-microphone array to emit a standard sound pulse with a duration of less than one millisecond. The system then analyzes the collected audio stream and, after detecting the direct pulse signal, identifies the first reflected echo with significant energy and accurately measures the time delay of the echo signal peak relative to the direct signal peak. , given that the speed of sound is , this delay Directly corresponds to the shortest acoustic reflection path between two points in the environment. In order to fundamentally ensure the purity of the original waveform segment template and avoid any contamination by early reflections, its predetermined length is automatically calculated and set to a fixed fraction of the first reflection delay, i.e. , which transforms the selection of segment length from a wide range to a deterministic operation with a unique solution based on field acoustic property measurements.
[0040] Accordingly, the generation of dynamic energy thresholds also follows a rigorous statistical procedure. During operation, the system continuously maintains a time period of Sliding time window, where It is set to 1 second, and the statistical characteristics of the audio stream energy in this window are calculated in real time, and the dynamic energy threshold It is continuously updated based on this, and its calculation formula is , in this formula, is the mean energy within the sliding window, is the standard deviation of the energy, and is a sensitivity coefficient, and its value is set to 3. Under the assumption that the signal energy obeys a Gaussian distribution, this setting can ensure the effective capture of energy mutation events exceeding three times the standard deviation, so that the trigger mechanism can accurately adapt to the non-steady-state fluctuations of the environmental background noise.
[0041] Furthermore, the calculation of the disorder index of the reverberation characteristic sequence is realized through a clear mathematical construction. When the system detects a sequence of reverberation characteristics caused by a source waveform segment, containing After obtaining a sequence of distortion-related peaks, the exact time point at which each peak reaches the microphone array is recorded to form a time series , the amplitudes of the corresponding peaks constitute an amplitude sequence , the algorithm first calculates the time interval sequence between adjacent peaks , and find the variance of this interval series , and secondly, by calculating the amplitude sequence The Spearman rank correlation coefficient of , to obtain an indicator to quantify its monotonically decreasing trend , disorder index It is finally given by the following linear weighted formula: , in, is a normalization constant whose value is determined during the offline calibration phase by analyzing the time interval variance of a large number of known background noise samples; The value range is , for an ideal reflection sequence, its value approaches 1; the weight coefficient and In this embodiment, both are set to 0.5 to equally balance the regularity of the time interval and the regularity of the amplitude attenuation. A sequence derived from physical reflection, Must be smaller and Approaching 1, resulting in a very low On the contrary, a random pseudo-peak sequence generated by diffuse noise will calculate a value close to 1. value.
[0042] The deep neural network mentioned in this invention serves as an auxiliary classification unit that works in conjunction with the aforementioned sound field physical information analysis core. This network utilizes the compact architecture of a one-dimensional convolutional neural network. Its sole input is the source waveform fragment captured by the algorithm. The network is trained offline using a training set consisting of two major categories of acoustic samples: typical transient effective sound sources from the target application scenario, and various types of background noise with stable waveform characteristics unique to that scenario. When the system is running online, each captured source waveform fragment is fed into the sliding cross-correlation computation channel and simultaneously fed into the solidified convolutional neural network for forward reasoning. The network outputs a confidence score that the fragment belongs to any of the pre-stored background noise categories. If this confidence score exceeds a preset threshold, the system uses this classification result as strong evidence and makes a joint decision with the calculated disorder index, thereby determining the acoustic event as background noise with greater certainty, effectively enhancing the system's immunity to complex interference.
[0043] Finally, the predetermined mutual correlation threshold and the predetermined disorder threshold in the algorithm are determined by an offline data-driven statistical calibration procedure. In the calibration process, the system first collects and stores a sufficient number of valid sound source event samples and background noise event samples that have been accurately labeled manually in the target environment. To determine the predetermined mutual correlation threshold, the algorithm calculates the pairwise mutual correlation coefficients between all valid sound source samples of different categories, and sets the threshold to a specific proportion of the observed minimum inter-class mutual correlation value, such as 90%, to ensure high sensitivity to novel sound sources; to determine the predetermined disorder threshold, the algorithm calculates the disorder index of all valid sound source samples and background noise samples respectively. , thus obtaining two statistical distributions representing two types of events respectively. The threshold is finally set at the optimal split point that minimizes the sum of the classification error rates of the two distributions.
[0044] Example 5: In this embodiment, the deterministic acquisition and application of a predetermined gain suppression factor follows an offline calibration procedure based on objective acoustic indicators. Specifically, the procedure first uses a standard transient sound source to excite a single acoustic event in a simulated environment with highly similar acoustic characteristics to the target application scenario. The system then completely collects the audio signal containing the direct sound and its subsequent reverberation. The algorithm then automatically identifies all audio signal segments determined to be reverberant. The system then traverses a candidate range of suppression factors with a preset step size. For each candidate suppression factor, the system multiplies the amplitude values of all sampling points of the identified reverberant signal segment by the factor to generate a processed audio version. Each processed version is then evaluated using an objective evaluation algorithm that can quantify the distortion of the target sound source signal and the level of reverberation interference suppression. Ultimately, the candidate factor that can maximize the suppression of reverberation while minimizing the distortion of the target sound source signal, i.e., the factor that obtains the optimal objective evaluation value, is determined and solidified as the predetermined gain suppression factor for that specific environment for subsequent online operation.
[0045] In other words, the solution for subtracting the energy of reverberation from the mixed audio signal is implemented through a clear signal processing process. The basis of this process is to convert the signal into the time-frequency domain to achieve accurate stripping in the energy dimension. Specifically, when one or more audio signal segments are identified as reverberation, the system first performs a short-time Fourier transform on the reverberation signal segment and calculates its average power spectrum to obtain a quantitative estimate of the energy of the reverberation in the frequency domain. Correspondingly, the system also performs a short-time Fourier transform on the entire mixed audio signal containing direct sound and reverberation frame by frame to obtain its power spectrum. In each time-frequency analysis frame, the system subtracts the reverberation obtained above from the power spectrum of the mixed signal. The average power spectrum is estimated to obtain the corrected power spectrum. To ensure the stability of the processing, the energy of any frequency bin that is lower than a preset non-negative lower limit after subtraction is set to the lower limit. Subsequently, the algorithm recombines the corrected power spectrum with the phase spectrum of the original mixed signal frame to construct a new complex spectrum. Finally, by performing an inverse short-time Fourier transform on the new complex spectrum, the signal is reconstructed from the time-frequency domain back to the time domain. By repeating this process for all audio frames and splicing them together, a purer target sound source signal with the reverberation component effectively separated in terms of energy is obtained. This procedure concretizes energy subtraction into a series of clearly defined spectral domain operations, ensuring the consistency and reproducibility of the separation effect.
[0046] Example 6: The judgment and tracking switching procedure for the moving sound source is specifically implemented by segmenting the captured source waveform segment and the subsequent 100 milliseconds of audio signal. The length of each audio frame is set to 40 milliseconds, the frame shift is 10 milliseconds, and the fundamental frequency of each frame is calculated using the YIN algorithm. Then, the fundamental frequency estimation values of 20 consecutive frames constitute an observation sequence, and the least squares linear regression analysis is performed on this sequence to obtain the slope of the fundamental frequency over time. , and continuously calculate 5 such observation sequences. If the slopes of all 5 sequences are The absolute value of each exceeds a slope threshold , and their signs remain consistent, the system determines that the sound source is in motion, where the slope threshold The value is based on the maximum radial motion speed of the sound source in the target scene. , through the Doppler effect formula Calculated, where is the typical fundamental frequency of the sound source, is the speed of sound; once it is determined to be a moving sound source, the system terminates the cross-correlation operation and switches to energy envelope tracking. This energy envelope is continuously generated by performing short-time root mean square energy calculation on the subsequent audio stream with a 20 millisecond frame length and a 5 millisecond frame shift. The energy envelope sequence is then used to calculate the arrival time difference of the dual-microphone array.
[0047] In addition, the application of deep neural networks in background noise recognition and the joint decision-making mechanism are as follows: each captured source waveform fragment is fed into the cross-correlation operation and simultaneously input into a one-dimensional convolutional neural network that has been trained offline. The network performs forward reasoning on the input fragment and outputs a posterior probability vector ,in is the total number of predefined background noise categories; the only condition for an acoustic event to be finally judged as background noise is the disorder index calculated based on the reverberation feature sequence Above its predetermined disorder threshold , and the maximum value of the posterior probability vector output by the aforementioned neural network Also above a predetermined confidence threshold ; The predetermined confidence threshold here Its value is determined during the offline validation phase of the network by selecting the confidence point corresponding to the preset false positive rate (FalsePositiveRate) of 1 / 10,000 on the receiver operating characteristic curve drawn on the validation set, thereby solidifying the decision preference of the classifier into the procedure.
[0048] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0049] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A multi-source sound field localization and separation algorithm based on deep neural network, characterized by: The algorithm comprises the following steps: Step a: obtaining an audio stream and continuously monitoring the energy of the audio stream; when the energy of the audio stream first exceeds a dynamic energy threshold dynamically adjusted based on the real-time average energy of the ambient background noise, intercepting an initial audio segment of a predetermined length from that moment and determining it as a source waveform segment, where the source waveform segment is the initial waveform characteristic of the sound source itself; Step b: Using the source waveform segment as a matching template, a sliding cross-correlation operation is performed on the audio stream following the source waveform segment to detect one or more distortion-related peaks that are delayed in the time domain and are highly correlated with the waveform features of the source waveform segment and have energy attenuation. The time delay and amplitude attenuation of the distortion-related peaks are information about the path and environment of sound propagation. Step c, determining the source waveform segment and the audio signal that follows it as the direct sound, and identifying the audio signal segment corresponding to one or more distortion-related peaks as the reverberation of the direct sound; performing sound source localization processing on the audio stream based on the identification results of the direct sound and the reverberation in the time domain; when an energy mutation is detected in the audio stream whose energy value exceeds the dynamic energy threshold and the cross-correlation with the source waveform segment currently in use is lower than the predetermined cross-correlation threshold, a new source waveform segment is generated, and the steps of detecting distortion-related peaks, determining direct sound and reverberation, and performing sound source localization processing are performed on the new source waveform segment in parallel.
2. The multi-source sound field localization and separation algorithm based on deep neural network according to claim 1 is characterized in that: After the distortion correlation peak detection step and before the sound source localization processing step, it also includes: continuously tracking the fundamental frequency of the original waveform segment and the audio signal that follows it to obtain the trajectory of the fundamental frequency changing with time; judging whether the fundamental frequency trajectory shows a continuous monotonic change; if it is judged as no, continuing to perform the sound source localization processing step; if it is judged as yes, terminating the sliding cross-correlation operation step and switching to a collaborative tracking method based on energy envelope and arrival time difference, wherein the continuous energy envelope of the sound source is continuously tracked, and the arrival time difference of the energy envelope is calculated using a dual-microphone array to obtain the direction of the moving sound source.
3. The multi-source sound field localization and separation algorithm based on deep neural network according to claim 1 is characterized in that: After the distortion correlation peak detection step and before the sound source localization processing step, the method further includes: constructing a reverberation feature sequence based on the occurrence time position and / or amplitude of one or more distortion correlation peaks in the time domain; calculating a disorder index of the reverberation feature sequence, where the disorder index is determined based on the variance of the time interval sequence and the monotonicity of the amplitude decay sequence in the reverberation feature sequence; executing the sound source localization processing step when the disorder index is lower than a predetermined disorder threshold; and terminating subsequent processing steps when the disorder index is higher than the predetermined disorder threshold, and identifying the current source waveform segment as background noise.
4. The multi-source sound field localization and separation algorithm based on deep neural network according to claim 3 is characterized in that: When the disorder index is higher than the predetermined disorder threshold and is identified as background noise, the current source waveform segment is marked as a background noise template. In subsequent audio stream monitoring, when the energy exceeds the dynamic energy threshold, the background noise template is preferentially used to match the energy mutation. If the match is successful, no subsequent processing is performed.
5. The multi-source sound field localization and separation algorithm based on deep neural network according to claim 1 is characterized in that: The predetermined length of the initial audio segment is five to ten milliseconds.
6. The multi-source sound field localization and separation algorithm based on deep neural network according to claim 1, characterized in that: The sliding cross-correlation operation is accelerated by fast Fourier transform.
7. The multi-source sound field localization and separation algorithm based on deep neural network according to claim 1 is characterized in that: The sound source localization process includes: when a dual-microphone array is used, the time delay difference between the first detected distortion correlation peak and the two microphones of the dual-microphone array is calculated. , calculate the azimuth of the reflecting surface Azimuth Calculated using the following formula: ,in, is the speed of sound, is the distance between the two microphones.
8. The multi-source sound field localization and separation algorithm based on deep neural network according to claim 1, characterized in that: The method further includes performing a sound source separation process on the audio stream, where the sound source separation process includes applying a predetermined gain suppression factor to perform gain suppression on an audio signal segment identified as reverberation.
9. The multi-source sound field localization and separation algorithm based on deep neural network according to claim 1, characterized in that: The method further includes performing a sound source separation process on the audio stream, where the sound source separation process includes subtracting the energy of the reverberation from the mixed audio signal.
Citation Information
Patent Citations
Hybrid speech processing method, electronic equipment and computer readable medium
CN120236599A
Acoustic emission positioning method, device and equipment
CN120254760A
Position detection system, transmission device, reception device, position detection method and position detection program
US20110116345A1