A multi-source sound field positioning and separation algorithm

By extracting original waveform segments from the audio stream and performing sliding cross-correlation calculations, combined with fundamental frequency tracking and disorder assessment, direct sound and reverberation separation in complex sound field environments is achieved. This solves the problems of inaccurate sound source localization and insufficient computing power in existing technologies, and realizes efficient multi-source sound field localization and separation.

CN120595237BActive Publication Date: 2025-10-24CHANGSHA HOTONE AUDIO
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511106087.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-10-24
Estimated Expiration
2045-08-08

AI Technical Summary

Technical Problem

Existing technologies struggle to dynamically separate direct and reflected sound in complex acoustic environments, leading to inaccurate sound source localization. Furthermore, the high computing power requirements exceed the capacity of edge devices, making real-time deployment impossible in mobile scenarios.

Method used

By acquiring the energy threshold trigger of the audio stream, the initial audio segment is extracted as the original waveform segment. The distortion correlation peak is detected by sliding cross-correlation operation. Combined with fundamental frequency tracking and disorder evaluation, the identification and separation of direct sound and reverberation are realized. The processing mode is dynamically switched, and the sound source location is calculated using a dual-microphone array.

Benefits of technology

Without the need for pre-set models and high-performance hardware, it achieves highly robust localization and separation of multi-source sound fields, reduces computational load, improves the accuracy and real-time performance of sound source identification, and reduces false triggering rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120595237B_ABST
    Figure CN120595237B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of speech or sound processing, and discloses a multi-source sound field positioning and separation algorithm, comprising: triggering source-born waveform segment interception through real-time monitoring of audio stream energy, taking it as a matching template to perform sliding cross-correlation operation with subsequent audio stream to detect distortion correlation peaks, and realizing sound source positioning and separation based on time domain identification results of direct sound and reverberation, wherein the present application analyzes the autocorrelation characteristics of sound waves, converts reverberation from interference signals into environmental messengers, enables the system to obtain the endogenous ability to analyze spatial information from reverberation, and simultaneously realizes autonomous and collaborative processing of transient and steady-state sound sources, thereby realizing high-robustness positioning in a complex sound field environment.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a multi-source sound field positioning and separation algorithm, belonging to the technical field of speech or sound processing. BACKGROUND

[0002] In the field of speech or sound processing, the positioning and separation of multi-source sound field is a key technology to realize industrial equipment anomaly monitoring, intelligent cockpit interaction and other scenes. The current mainstream scheme generally relies on deep neural networks to build an end-to-end acoustic mapping model, and learns the correlation between sound source features and spatial positions through massive data. Although this method can capture some acoustic laws, it has two fundamental limitations: first, it simplifies the sound field as a linear superposition signal, ignoring the nonlinear interference effect caused by the reflection of sound waves on rigid boundaries such as equipment shells and walls, resulting in feature confusion when sudden sound sources (such as metal collisions) and steady-state noise (such as engine humming) are coupled; second, to solve the problem of insufficient environmental adaptability, the network depth is increased, forcing the algorithm power demand to far exceed the carrying capacity of edge devices, making it difficult to deploy in real time in mobile scenarios such as industrial inspection robots.

[0003] Taking a forklift collision detection scene as an example, the existing technology needs to pre-set a sound field environment geometric model to suppress reverberation interference. However, when sound waves pass through the shelf partition area, the difference in acoustic impedance between metal and cloth materials causes a calculation deviation of more than 30% in the reflection path, forcing the system to frequently switch to a low-precision direct sound separation mode. The industry tries to introduce high-precision motion sensors to compensate for the Doppler shift, but the addition of hardware costs and real-time calibration burdens further exacerbate the contradiction between system complexity and power consumption.

[0004] The deeper technical contradiction is that the existing paradigm regards reverberation as noise that needs to be eliminated, rather than a physical messenger that contains spatial information. This cognitive bias traps the system in a passive cycle of data fitting the environment - forced to stack network parameters to improve accuracy, but unable to fundamentally solve the robustness problem of sound source positioning in dynamic environments due to the neglect of universal laws of acoustic boundary effects. Specifically, the existing technology mainly has the following core bottlenecks: 1. Relies on pre-set CAD models, unable to respond to changes in scene geometry in real time (such as moving device displacement); 2. Separates physical variables such as material acoustic impedance difference and sound source motion state from signal processing, resulting in insufficient model generalization; 3. Complex models and edge hardware capabilities are severely mismatched, restricting the technology's ability to reach the masses. Therefore, how to analyze and reconstruct the environmental cognition mechanism through the autocorrelation characteristics of sound waves to achieve dynamic decoupling of multi-source sound fields without high computing power and pre-set models has become a technical problem to be solved by the present application. SUMMARY

[0005] The present application provides a multi-source sound field positioning and separation algorithm, which mainly aims to solve the problem of dynamic separation of direct sound and reflected sound in complex sound field environments.

[0006] To achieve the above object, the application provides a multi-source sound field positioning and separation algorithm, which comprises the following steps:

[0007] Step a: obtaining an audio stream and continuously monitoring the energy of the audio stream; when the energy of the audio stream first exceeds a dynamic energy threshold dynamically adjusted according to the real-time average energy of the environmental background noise, an initial audio segment with a predetermined length from the moment is intercepted, and the initial audio segment is determined as a source waveform segment, which is the initial waveform characteristic of the sound source itself;

[0008] Step b: taking the source waveform segment as a matching template, performing a sliding cross-correlation operation on the audio stream after the source waveform segment to detect one or more distortion-related peaks with high correlation to the waveform characteristics of the source waveform segment and energy attenuation in the time domain, the time delay and amplitude attenuation of the distortion-related peaks being the path and environmental information of sound propagation;

[0009] Step c: determining the source waveform segment and the audio signal immediately following the source waveform segment as direct sound, and identifying the audio signal segment corresponding to the one or more distortion-related peaks as reverberation of the direct sound; performing sound source positioning processing on the audio stream according to the identification results of the direct sound and the reverberation in the time domain; when an energy mutation with an energy value exceeding the dynamic energy threshold and a cross-correlation lower than a predetermined cross-correlation threshold with the source waveform segment currently being used is detected in the audio stream, a new source waveform segment is generated, and the steps of detecting distortion-related peaks, determining direct sound and reverberation, and performing sound source positioning processing are performed in parallel on the new source waveform segment, thereby realizing parallel positioning and separation of multiple sound sources.

[0010] Preferably, after the step of detecting distortion-related peaks and before the step of sound source positioning processing, the method further comprises: continuously tracking the fundamental frequency of the source waveform segment and the audio signal immediately following the source waveform segment to obtain a trajectory of the fundamental frequency changing with time; judging whether the fundamental frequency trajectory presents a continuous monotonic change; if not, the step of sound source positioning processing is continued; if yes, the step of performing a sliding cross-correlation operation is suspended, and a cooperative tracking method based on energy envelope and time difference of arrival is switched to, wherein the continuous energy envelope of the sound source is continuously tracked, and the time difference of arrival of the energy envelope is calculated by using a double-microphone array, thereby obtaining the direction of the moving sound source.

[0011] Preferably, after the step of detecting the distortion-related peaks and before the step of sound source localization processing, the method further comprises: based on the time position and / or amplitude of the one or more distortion-related peaks in the time domain, constructing a reverberation feature sequence; calculating a disorder degree index of the reverberation feature sequence, the disorder degree index being determined according to the variance of the time interval sequence and the monotonicity degree of the amplitude decay sequence in the reverberation feature sequence; when the disorder degree index is lower than a predetermined disorder degree threshold, performing the step of sound source localization processing; when the disorder degree index is higher than the predetermined disorder degree threshold, then suspending the subsequent processing steps, and identifying the current source-generated waveform segment as background noise.

[0012] Preferably, when the disorder degree index is higher than the predetermined disorder degree threshold and is identified as background noise, the current source-generated waveform segment is marked as a background noise template, and in subsequent audio stream monitoring, when the energy exceeds the dynamic energy threshold, the background noise template is preferentially used for matching with the energy mutation, and if the matching is successful, subsequent processing is not performed.

[0013] Preferably, the predetermined length of the initial audio segment is five milliseconds to ten milliseconds.

[0014] Preferably, the sliding cross-correlation operation is accelerated by fast Fourier transform.

[0015] Preferably, the sound source localization processing comprises: when a double-microphone array is used, according to the time delay difference between the arrival of the first detected distortion-related peak at the two microphones of the double-microphone array , calculating the azimuth angle of the reflecting surface Azimuth angle According to the following formula: , wherein, is the speed of sound, is the distance between the two microphones.

[0016] Preferably, the method further comprises performing sound source separation processing on the audio stream, the sound source separation processing comprising: applying a predetermined gain suppression factor to the audio signal segment identified as reverberation for gain suppression.

[0017] Preferably, the method further comprises performing sound source separation processing on the audio stream, the sound source separation processing comprising: subtracting the energy of the reverberation from the mixed audio signal.

[0018] Preferably, the method is performed by a microcontroller, the microcontroller being configured to perform the energy detection operation, the audio slicing operation, and the one-dimensional cross-correlation operation.

[0019] Compared with the prior art, the present application has the following beneficial effects:

[0020] 1. By using the initial waveform of the sound source as an autocorrelation template, the system acquires for the first time the intrinsic ability to parse spatial information from reverberation. When sound waves propagate in a complex environment, the temporal correlation between the direct sound and its distortion-related peaks enables the system to directly utilize the path delay information of the reflected sound to naturally distinguish the sound source itself from environmental reflectors. This cognitive leap, transforming reverberation from an interference signal to an environmental messenger, avoids the complex calculations of traditional solutions to combat reverberation, and shifts sound field understanding from passive decomposition to active perception.

[0021] 2. The combination of a native waveform capture mechanism triggered by dynamic energy thresholds and multi-event parallel processing logic enables the system to automatically distinguish between sound source events of different natures through waveform correlation in scenarios where sudden sound sources such as metal collisions coexist with steady-state noise such as machine roars. The system no longer requires preset models for specific sound source types. Instead, it spontaneously forms multi-source processing channels based on the intrinsic time-domain fingerprint characteristics of sound waves, which in principle resolves the performance conflicts of traditional models when heterogeneous sound sources are mixed. When the continuous drift of the fundamental frequency trajectory reveals the motion state of the sound source, the system dynamically switches to energy envelope tracking mode. This autonomous evolution of processing logic eliminates the Doppler effect of the moving sound source, which no longer causes positioning failure, but instead transforms it into a natural indicator of motion detection. By reusing the arrival time difference calculation of the two microphones, the system maintains the continuity of sound source direction tracking while avoiding complex motion compensation models, achieving a natural unification of static analysis and dynamic tracking.

[0022] 3. The disorder evaluation mechanism of the reverberation peak sequence enables the system to identify the acoustic background. When a disordered correlation peak caused by steady-state noise is detected, the system automatically terminates invalid calculations and marks the current waveform as a noise template. This mechanism makes lightweight judgments on information, making the spectral characteristics of the background noise an immune marker for the system's self-learning, significantly reducing the false trigger rate while maintaining sensitivity to weak valid signals. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Fig. 1 Schematic diagram of the sound wave propagation path and reflection characteristics of the present invention;

[0024] Fig. 2 This is a timing diagram of the background noise template recognition and disorder determination process of the present invention;

[0025] Fig. 3 This is a comparison chart of the calculated values ​​of the disorder index under different sound source types of the present invention.

[0026] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0027] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0028] The present invention discloses a multi-source sound field localization and separation algorithm, whose overall architecture and operation process mainly consist of core stages such as parallel capture of source waveforms, collaborative identification of sound source states, and temporal decoupling of sound field information. These stages work together to transform reverberation signals in complex environments from interference sources in traditional cognition into physical messengers carrying spatial information, thereby achieving highly robust localization and separation of multiple static and dynamic sound sources without the need for preset environmental models and high-computing hardware.

[0029] Specifically, the typical application scenario is an intelligent inspection robot or automatic forklift deployed in a complex indoor environment such as a factory or warehouse. The core processor of the system is usually an embedded microcontroller, which continuously performs real-time acquisition and analysis of the audio stream. The microcontroller first continuously monitors the environmental background noise, and dynamically generates and adjusts a dynamic energy threshold by calculating the real-time average energy within a time window. The threshold is set to be slightly higher than the steady-state noise level of the current environment. This ensures that the system has a sensitive perception of energy mutations, while effectively avoiding false triggering caused by meaningless fluctuations in background noise. When a transient acoustic event occurs in the scene, such as the impact sound of a metal tool falling from a height, the energy of the sound signal will instantaneously and significantly exceed the aforementioned dynamic energy threshold. At this time, the system is immediately triggered, and from the moment the energy mutation occurs, the system is intercepted. An initial audio clip of a predetermined length is taken; the length of the clip is set between five and ten milliseconds. The selection of this length is based on a key technical trade-off: the duration of five to ten milliseconds is sufficient to fully capture the unique initial waveform fingerprint of a sound source such as a metal impact sound, but short enough to ensure with a high probability that the intercepted clip contains only direct sound without any reflection, thereby ensuring the purity of the template at the source; the intercepted initial audio clip is then determined by the system as a source waveform clip, which serves as an acoustic fingerprint template representing the physical characteristics of the sound source itself. After successfully capturing the source waveform clip, the system immediately enters the sliding cross-correlation operation stage, in which the source waveform clip is used as a matching template to perform continuous sliding matching calculations on the real-time audio stream that follows it; to ensure that the operation can be efficiently executed on a resource-constrained microcontroller, the sliding cross-correlation operation is performed through a fast Fourier transform The acceleration converts the computationally intensive convolution operation in the time domain into the more efficient multiplication operation in the frequency domain, greatly reducing the computational load; the core goal of the operation is to detect one or more delay occurrences in the time domain, which are highly correlated with the source waveform segment in waveform characteristics, but have some attenuation in energy; these distortion-related peaks are essentially reverberations formed by the reflection of direct sound through different physical paths, and the time delay of these peaks relative to the direct sound directly maps the length difference of the sound propagation path, while the amplitude attenuation reflects the acoustic absorption characteristics of the reflection interface and the energy loss caused by the propagation distance, therefore, the sequence of these distortion-related peaks contains rich information about the geometric structure of the sound field environment.

[0030] Based on the above operation results, the system identifies the source waveform segment as the matching template and its first occurrence position in the audio stream as the direct sound, and identifies all subsequent detected distortion-related peaks corresponding to the audio signal segment as the reverberation of the direct sound; by analyzing the time delay difference between the direct sound and the first strongest distortion-related peak, which is usually the reflected sound produced by the nearest single reflection surface, between the dual-microphone array deployed on the robot, the system can perform preliminary sound source positioning processing; specifically, if the fixed distance between the two microphones of the dual-microphone array is , the sound speed is , and the system measures the time delay difference of the first detected distortion-related peak to the two microphones as , then the system can calculate the azimuth of the physical reflection surface producing the reflection relative to the robot's own coordinate system according to the following formula : By solving the azimuth of one or more such reflections, the system can inversely calculate the possible position area of the sound source, achieving high-precision positioning; in addition to this positioning process, the technical solution also designs a parallel processing and state switching mechanism to deal with complex acoustic scenes; when the system is processing a sound source event such as a metal impact sound, if another energy mutation with energy value exceeding a dynamic energy threshold appears in the audio stream, for example, a worker nearby makes a shout, the system will first perform cross-correlation calculation on the new energy mutation waveform and the currently used source waveform segment; since the waveform characteristics of the shout and the metal impact sound are completely different, their cross-correlation will be significantly lower than a preset cross-correlation threshold, based on this judgment, the system recognizes it as a new, independent sound source event, so the system generates a completely new source waveform segment for the shout, and opens up a parallel processing channel for it, independently performing the subsequent steps of detecting distortion-related peaks, determining direct sound and reverberation, and performing sound source positioning processing, thereby achieving parallel positioning and separation of multiple sound sources of different sources and different natures.

[0031] Further, to handle the special case of moving sound sources, such as the engine sound of a forklift truck that is continuously approaching or moving away, the system introduces a judgment logic for the motion state of the sound source after the detection of the distortion-related peaks. The system performs continuous fundamental frequency tracking on the source waveform segment and the audio signal immediately following it to obtain the trajectory of the fundamental frequency over time. If the fundamental frequency trajectory shows a clear and continuous monotonic change, whether it is a monotonic increase or a monotonic decrease, it constitutes clear physical evidence of the Doppler effect, indicating that the sound source is moving radially relative to the robot. Once it is confirmed that the sound source is in motion, the system determines that the positioning method based on reverberation may lose accuracy due to the rapid changes in the position of the sound source. Therefore, the system will suspend the subsequent sliding cross-correlation operation step and dynamically switch to a collaborative tracking method based on energy envelope and time difference of arrival. In this method, the system instead continuously tracks the continuous energy envelope of the moving sound source and calculates the time difference of arrival of the energy envelope between the two microphones in real time, thereby continuously and stably tracking the direction of the moving sound source, achieving intelligent and seamless switching between the two modes of static positioning and dynamic tracking.

[0032] In addition, to enhance the system's ability to recognize and exempt environmental background noise, the system also includes an out-of-order degree evaluation mechanism based on the reverberation feature sequence. After detecting a series of distortion-related peaks, the system will form a reverberation feature sequence based on the time positions and amplitudes of these peaks in the time domain. Then, the system will calculate the out-of-order degree index of the sequence, which takes into account two dimensions: the variance of the time interval sequence and the monotonicity degree of the amplitude decay sequence. A reverberation sequence composed of clear physical reflections will typically exhibit strong physical regularity in the time interval and energy decay, resulting in a low out-of-order degree index. On the other hand, if a sound source is essentially diffuse and continuous background noise, such as a large fan or air conditioning system, its autocorrelation operation may produce a series of pseudo-peaks that appear to be correlated but have chaotic timing and energy decay regularity, resulting in an out-of-order degree index higher than a predetermined out-of-order degree threshold. When this occurs, the system determines that the current source waveform segment is not derived from a clear acoustic event, but from background noise, and thus suspends the subsequent positioning and separation processing steps. Thereafter, to further improve system efficiency, the system will mark this source waveform segment identified as background noise as a background noise template. In subsequent audio stream monitoring, when a new energy mutation exceeds the threshold, the system will preferentially use the stored background noise template to match the mutation. If the match is successful, it is directly determined as known background noise and ignored, without initiating any subsequent processing, thereby achieving adaptive immunity to specific background noise and significantly reducing the false trigger rate and unnecessary computational overhead of the system.

[0033] Finally, after the positioning of the sound source is completed, the technical solution also provides a sound source separation processing function to improve the clarity of the target sound source. One separation method is to apply a predetermined gain suppression factor to all audio signal segments identified as reverberation, i.e., significantly reducing the volume without completely eliminating it. Another more direct method is to estimate and confirm the total energy of the reverberation part, and then subtract this part of the reverberation energy from the mixed overall audio signal. In this way, the interference of environmental reflection on the main sound source can be effectively suppressed, and a purer direct sound signal can be extracted, ultimately improving the application performance of the entire system.

[0034] Embodiment 1: The technical solution of the present application in a continuously running full-automatic stereoscopic warehouse environment, the specific operation mode is as follows: the warehouse is deployed with a high-speed running conveying and sorting system, and a driverless forklift travels between metal shelves for operation, when a cargo packaging box falls from the conveying belt and hits the ground, the short impact sound wave energy immediately exceeds the dynamic energy threshold dynamically adjusted according to the real-time average energy of the environmental background noise, the system immediately intercepts a predetermined initial audio segment of five to ten milliseconds from this moment as the source waveform segment of the impact event; the system immediately takes this source waveform segment as a matching template to perform a sliding cross-correlation operation based on fast Fourier transform acceleration on the audio stream that follows, and detects a series of distortion-related peaks in time domain that are highly correlated with the source waveform characteristics but have decaying energy, by converting the reverberation signal from the interference to be eliminated into an environmental messenger containing spatial information, the system fundamentally redefines the technical paradigm of sound field perception without relying on any pre-set warehouse geometric model, and changes the understanding of the environment from passive model fitting to active analysis based on the autocorrelation characteristics of sound waves, while processing the above impact event, the engine roar of a driverless forklift accelerating is processed by the system as an independent acoustic event, because its waveform characteristics have low cross-correlation with the source waveform segment of the impact sound, the system generates a new source waveform segment for it; after continuous fundamental frequency tracking of the new segment, the system identifies that its fundamental frequency trajectory presents a continuous monotonic change, which reveals that the sound source is in motion and produces a Doppler effect, at this time, a synergistic mechanism is triggered: the sound source motion state obtained by the fundamental frequency tracking provides a decisive prior condition for subsequent processing, and the system stops the sliding cross-correlation operation for the sound source and switches to a synergistic tracking method based on energy envelope and time difference of arrival, using a double microphone array to calculate the time delay difference of energy envelope arriving at two microphones to obtain the direction of the sound source. The autonomous switching of the two modes of static analysis and dynamic tracking realizes the matching of processing logic and physical state of the sound source, greatly reduces the operation load, and ensures the continuity and accuracy of tracking different nature sound sources, in this process, the ventilation system at the top of the warehouse continuously emits stable low-frequency humming, and the energy fluctuation may also trigger the energy threshold for a short time, the system will also generate a source waveform segment for it and start cross-correlation detection, however, the distortion-related peak sequence generated by it shows high randomness in time domain distribution and amplitude decay;The system calculates the disorder index of the reverberation characteristic sequence—determined by the variance of the time interval sequence and the monotonicity of the amplitude decay sequence—and finds that its value exceeds a predetermined threshold. This mechanism physically resolves the inherent contradiction between increasing sensitivity and suppressing false triggers in the industry. The system does not rely on a voiceprint database, but instead, based on the inherent laws of sound wave propagation, clearly identifies the waveform as background noise and marks it as a background noise template. In subsequent monitoring, any energy mutations matching this template are directly ignored, thus achieving effective immunity to steady-state environmental noise without compromising the ability to capture true transient events.

[0035] The overall architecture of the algorithm, by internalizing the laws of acoustic physics into an analysis of signal timing and intrinsic correlation, makes the geometric structure of the sound field environment and the motion state of the sound source no longer external variables that require preset models to passively adapt, but endogenous information that can be perceived in real time through the signal's own evolutionary characteristics. This design avoids the traditional technical path's reliance on massive labeled data, complex neural network models, and high-computing-power hardware, allowing highly robust multi-source sound field positioning and separation functions to be implemented on a microcontroller that only needs to perform basic operations such as energy detection, audio slicing, and one-dimensional cross-correlation calculations.

[0036] Example 2: To quantitatively verify the core capability of the present invention in distinguishing and locating transient impact sound sources from steady-state background noise in complex sound fields, this example aims to accurately test the algorithm's ability to effectively capture burst signals, spatially locate them using their reverberation information, and autonomously identify and suppress coexisting steady-state noise, without a preset environmental model. This demonstrates the technical effectiveness of the present invention in converting reverberation from an interference signal into an environmental messenger and achieving autonomous collaborative processing of heterogeneous sound sources. To this end, a reproducible physical test platform simulating an industrial environment was constructed. The platform was built in a test room measuring 5 meters by 5 meters by 3 meters. Metal plates and sound-absorbing cotton were arranged alternately on the walls to simulate the highly complex acoustic reflection environment of a factory workshop where shelves and open spaces coexist. The acoustic sensing core was a dual-microphone array with a spacing of 15 cm, placed in a corner of the room as a signal receiver. The sound source consisted of two parts: one deployed at the indoor coordinates 4.0 and 3.5. An industrial-grade loudspeaker continuously plays a recording of a steady-state forklift engine noise with an average sound pressure level of 75 decibels. An electrically controlled air cannon, capable of generating instantaneous, high-intensity shocks with a peak sound pressure level of 110 decibels, can be triggered from anywhere in the room to simulate sudden events such as falling cargo or equipment collisions. A key parameter in the experiment is the predetermined length of the source waveform segment, which must strike a technical balance between ensuring waveform uniqueness and avoiding early reverberation contamination. The selection of this parameter directly affects the accuracy of the subsequent sliding cross-correlation operation and is a prerequisite for determining whether the system can successfully decouple the direct sound signature from the signal. The core of this technical trade-off is that if the segment is too short, it may not contain sufficient waveform detail to distinguish it from other acoustic events, resulting in a lower signal-to-noise ratio for the correlation analysis. Conversely, if the segment is too long, it is very likely to include primary reflections from the reflective surface closest to the sound source, thereby contaminating the source waveform used as the matching template and losing its purity as a representation of the direct sound. Therefore, the decision rule for this length is modeled as follows: its duration should be significantly longer than the energy rise time of the target transient sound source to ensure that the core features are captured, and at the same time it must be less than the first reflection delay determined by the minimum reflection path in the room to ensure the purity of the template. In this test environment, the main energy pulse of the air cannon impact sound is completed within 2 milliseconds, and the first reflection sound delay caused by the sound source being at least 1 meter away from the nearest wall is approximately Based on this decision model, a non-restrictive sample length of 8 milliseconds is set. This value not only completely covers the core waveform of the impact sound, but also effectively avoids the interference of early reflections, laying a solid foundation for subsequent autocorrelation analysis.

[0037] After the test started, the system first ran in an environment with only the forklift engine steady noise. The algorithm monitored that the audio stream energy continuously and steadily exceeded the energy threshold dynamically generated according to the background silence level, and then intercepted the first 8 ms audio segment as the initial source waveform segment. The system used this as a template to perform a sliding cross-correlation operation on the subsequent audio stream. At this time, the key phenomenon observed was that the correlation peak sequence presented a high density, low attenuation, and disordered morphology, that is, a large number of correlation peaks appeared densely, the variance of their time intervals was extremely large, and the amplitude attenuation sequence did not have clear monotonicity. The system quantitatively calculated this according to the reverberation characteristic sequence disorder degree evaluation mechanism of the present application, as shown in Table 1.

[0038] Table 1: Different sound source characteristics and system determination results table.

[0039] As shown in the above table data, the disorder degree index calculation value of the waveform segment triggered by the steady noise is 0.91, which is significantly higher than the disorder degree threshold of 0.8 preset by the system. The internal mechanism of this data result is that the reverberation sound field formed by the steady noise in a complex environment is a random process of continuous superposition and energy dispersion. Its reflection path presents a chaotic arrival in the time domain, lacking the clear time sequence structure from the direct sound to the various reflected sounds of the transient impact sound source. Therefore, the algorithm successfully identifies and labels this source waveform segment as a background noise template according to this high disorder degree index, and stops the invalid calculation of positioning it, thereby avoiding the waste of computing power, which directly verifies the ability of the present application to identify and exempt background noise using the reverberation disorder degree evaluation mechanism. Subsequently, in the background of continuous engine noise, the air cannon located at coordinates 2.5, 2.0 was triggered. The system monitored a sharp energy mutation with an energy value far exceeding the dynamic threshold and a cross-correlation with the stored background noise template lower than the preset threshold 0.3. According to the multi-event parallel processing logic of the present application, the system immediately determines that this is a new and independent acoustic event, and generates a completely new 8 ms source waveform segment for it. The subsequent sliding cross-correlation operation presents a completely different phenomenon from before: a clear, highest-energy direct sound signal, and a series of distorted correlation peaks with regular energy decay appearing in time delay. The system accurately records the arrival time of these key signals using a dual-microphone array, as shown in Table 2.

[0040] Table 2: Arrival time difference table of different signal types.

[0041] The system uses the first reflection peak as the signal to analyze spatial information, and uses its time delay difference microseconds to the dual-microphone array to calculate the azimuth angle of the reflection surface according to the formula , the calculation process is The calculation results are consistent with the air cannon and the nearest reflecting surface, which is located at This process verifies the core mechanism of the present invention: by using the sound source's own waveform as a template for autocorrelation matching, the complex reverberation signal is successfully decoupled into a distorted correlation peak sequence containing propagation path and delay information, thereby converting the reverberation that is regarded as interference in traditional methods into key physical information for highly robust sound source localization.

[0042] Example 3: This example combines Figs. 1 to 3 , a multi-source sound field localization and separation algorithm is implemented. Fig. 1 As shown in the figure, the horizontal axis is time (ms) and the vertical axis is amplitude. By marking and classifying the waveforms of multiple key signals in the process of sound wave propagation, the amplitude attenuation and time delay characteristics of the sound wave signal under different propagation paths are reflected. The first group of waveforms on the far left is represented by a thick solid line and is marked as the source waveform. It has the largest amplitude and represents the direct sound signal initially emitted by the sound source, corresponding to the line description of the direct sound in the legend; followed by As a mark, a group of waveforms represented by dotted lines appear, which are the first reflection signals, marked as the first reflection, representing the signal response after the sound wave is reflected by an obstacle once; followed by As intervals, a dotted line waveform appears, marked as secondary reflection, indicating the propagation signal of the sound wave after two reflections; Reflected waveforms continue to appear at intervals, with their amplitudes gradually weakening, indicating the gradual attenuation of signal energy due to multiple reflections.

[0043] like Fig. 2 As shown, first, the main control module initiates an instruction to the reverberation analysis module to request analysis of the time-varying correlation peak sequence. After completing the analysis, the reverberation analysis module returns the reverberation feature sequence. Then the main control module requests to calculate the disorder based on the feature sequence. This request is sent to the disorder evaluation module. The disorder evaluation module evaluates the disorder of the waveform by calculating the time interval variance and the amplitude decay monotonicity, and returns the disorder index value to the main control module. The main control module determines whether the index is higher than the predetermined threshold based on the return value. If the result satisfies [the index is higher than the threshold], the main control module determines the current waveform as background noise, and continues to mark and store it as a background noise template. After the operation is completed, the main control module further confirms that the template has been stored, and executes the termination of subsequent positioning and separation processing, terminating the subsequent sound source positioning and signal separation process of the background noise waveform.

[0044] like Fig. 3As shown, the horizontal axis is the calculated value, and the vertical axis is the monotonicity of the amplitude attenuation sequence, the variance of the time interval sequence, and the calculated disorder index. The three parameters are used to evaluate the attenuation regularity, time consistency, and overall disorder of the sound wave characteristics, respectively. The figure compares two types of sound sources, namely air cannon impact and engine steady-state noise, which are marked with different shapes in the legend. The circle represents the air cannon impact and the square represents the engine steady-state noise. It can be seen from the data in the figure that the calculated values ​​of the air cannon impact in terms of the monotonicity of the amplitude attenuation sequence, the variance of the time interval sequence, and the calculated disorder index are significantly lower than those of the engine steady-state noise, especially in the dimension of the calculated disorder index. The corresponding calculated value of the air cannon impact is about 0.1, while the corresponding value of the engine steady-state noise is close to 0.9.

[0045] Example 4: In this example, when the system is first deployed in a specific working environment, a single environmental self-calibration process is performed to set the predetermined length of the source waveform segment. After the process is started, the system drives a miniature acoustic exciter configured near its dual-microphone array to emit a standard sound pulse with a duration of less than one millisecond. The system then analyzes the collected audio stream and, after detecting the direct pulse signal, identifies the first reflected echo with significant energy and accurately measures the time delay of the echo signal peak relative to the direct signal peak. , given that the speed of sound is , this delay Directly corresponds to the shortest acoustic reflection path between two points in the environment. In order to fundamentally ensure the purity of the original waveform segment template and avoid any contamination by early reflections, its predetermined length is automatically calculated and set to a fixed fraction of the first reflection delay, i.e. , which transforms the selection of segment length from a wide range to a deterministic operation with a unique solution based on field acoustic property measurements.

[0046] Accordingly, the generation of dynamic energy thresholds also follows a rigorous statistical procedure. During operation, the system continuously maintains a time period of Sliding time window, where It is set to 1 second, and the statistical characteristics of the audio stream energy in this window are calculated in real time, and the dynamic energy threshold It is continuously updated based on this, and its calculation formula is , in this formula, is the mean energy within the sliding window, is the standard deviation of the energy, and is a sensitivity coefficient, and its value is set to 3. Under the assumption that the signal energy obeys a Gaussian distribution, this setting can ensure the effective capture of energy mutation events exceeding three times the standard deviation, so that the trigger mechanism can accurately adapt to the non-steady-state fluctuations of the environmental background noise.

[0047] Further, for the calculation of the disorder index of the reverberation feature sequence, a clear mathematical construction is implemented. When the system detects a sequence containing distortion-related peaks triggered by a source-born waveform segment, it records the exact time points of each peak arriving at the microphone array to form a time sequence , whose amplitudes corresponding to each peak form an amplitude sequence . The algorithm first calculates the time interval sequence between adjacent peaks , and obtains the variance of this interval sequence . Secondly, by calculating the Spearman rank correlation coefficient of the amplitude sequence , an index quantifying its monotonic decreasing trend is obtained . The disorder index is finally given by the following linear weighting formula:

[0048] ,

[0049] where is a normalization constant, whose value is determined by analyzing the time interval variance of a large number of known background noise samples in the offline calibration stage; The value range of is , and its value tends to 1 for an ideal reflection sequence; the weight coefficients and are both set to 0.5 in this embodiment, to equally weigh the regularity of time interval and the regularity of amplitude decay. A sequence resulting from physical reflection must have a smaller and tends to 1, thus producing a very low value; on the contrary, a random pseudo-peak sequence generated by diffuse noise will calculate a value close to 1.

[0050] The deep neural network mentioned in the present application is an auxiliary classification unit that works in cooperation with the aforementioned sound field physical information analysis core. The network adopts a compact architecture of one-dimensional convolutional neural network, and the only input of the network is the source waveform segment captured by the algorithm. The training of the network is completed in an offline state, and the training set used contains two categories of acoustic samples: one is various typical transient effective sound sources collected from the target application scenario, and the other is various background noises with stable waveform characteristics specific to the scene. During the online operation of the system, each captured source waveform segment is simultaneously sent into the convolutional neural network for forward inference when it is sent into the sliding cross-correlation operation channel. The network outputs the confidence degree of the segment belonging to any pre-stored background noise category. If the confidence degree exceeds a pre-set threshold, the system regards this classification result as strong evidence and makes a joint decision with the calculation result of the disorder degree index, so as to determine the acoustic event as background noise with higher certainty, thereby effectively enhancing the exemption ability to complex interference.

[0051] Finally, the values of the predetermined cross-correlation threshold and the predetermined disorder degree threshold in the algorithm are determined by an offline data-driven statistical calibration procedure. In the calibration process, the system first collects and stores sufficient samples of effective sound source events and background noise events in the target environment, which have been accurately labeled by humans. To determine the predetermined cross-correlation threshold, the algorithm calculates the pairwise cross-correlation coefficients between all different categories of effective sound source samples, and sets the threshold as a certain proportion, such as 90%, of the observed minimum cross-correlation coefficient value between categories, to ensure high sensitivity to new and unusual sound sources. To determine the predetermined disorder degree threshold, the algorithm calculates the disorder degree index of all effective sound source samples and background noise samples respectively , thereby obtaining two statistical distributions respectively representing the two categories of events. The threshold is finally set at the optimal segmentation point that minimizes the sum of classification error rates of the two distributions.

[0052] In this embodiment, the determination of the predetermined gain suppression factor is achieved in a deterministic manner, following an offline calibration procedure based on objective acoustic metrics. Specifically, the procedure first triggers a single acoustic event in a simulated environment that is highly similar to the target application scenario or acoustic characteristics, using a standard transient sound source, and the system fully captures the audio signals containing the direct sound and its subsequent reverberation. The algorithm then automatically identifies all audio signal segments that are determined to be reverberation. Subsequently, the system traverses a candidate interval of suppression factors with a preset step size. For each candidate suppression factor, the system generates a processed audio version by multiplying the amplitude values of all sampling points of the identified reverberation signal segments by the factor. Then, an objective evaluation algorithm that can quantify the distortion of the target sound source signal and the level of reverberation suppression is used to evaluate each processed version. Finally, the candidate factor that can suppress the reverberation to the greatest extent while minimizing the distortion of the target sound source signal, i.e., the factor that obtains the optimal objective evaluation value, is determined and fixed as the predetermined gain suppression factor for the specific environment, which is used for subsequent online operation.

[0053] Furthermore, the scheme for subtracting the energy of the reverberation from the mixed audio signal is implemented through a clear signal processing procedure. The basis of this procedure is to perform time-frequency domain conversion on the signal to achieve accurate stripping in the energy dimension. Specifically, after one or more audio signal segments are identified as reverberation, the system first performs short-time Fourier transform on the reverberation signal segment and calculates its average power spectrum to obtain an energy quantitative estimate of the reverberation in the frequency domain. Accordingly, the system also performs short-time Fourier transform on the entire mixed audio signal containing the direct sound and the reverberation to obtain its power spectrum. In each time-frequency analysis frame, the system subtracts the aforementioned obtained average power spectrum estimate of the reverberation from the power spectrum of the mixed signal to obtain a modified power spectrum. To ensure the stability of the processing, any frequency bin energy that is lower than a preset non-negative bottom limit value after subtraction is set to the bottom limit value. Subsequently, the algorithm recombines the modified power spectrum with the phase spectrum of the original mixed signal frame to construct a new complex spectrum. Finally, the signal is reconstructed from the time-frequency domain to the time domain by performing inverse short-time Fourier transform on the new complex spectrum. By repeating this process for all audio frames and splicing them together, a purer target sound source signal with the reverberation component effectively separated in energy is obtained. This procedure concretizes the energy subtraction into a series of clearly defined spectral domain operations, ensuring the consistency and reproducibility of the separation effect.

[0054] In this embodiment, the determination of the predetermined gain suppression factor is achieved in a deterministic manner, following an offline calibration procedure based on objective acoustic metrics. Specifically, the procedure first triggers a single acoustic event in a simulated environment that is highly similar to the target application scenario or acoustic characteristics, using a standard transient sound source, and the system fully captures the audio signals containing the direct sound and its subsequent reverberation. The algorithm then automatically identifies all audio signal segments that are determined to be reverberation. Subsequently, the system traverses a candidate interval of suppression factors with a preset step size. For each candidate suppression factor, the system generates a processed audio version by multiplying the amplitude values of all sampling points of the identified reverberation signal segments by the factor. Then, an objective evaluation algorithm that can quantify the distortion of the target sound source signal and the level of reverberation suppression is used to evaluate each processed version. Finally, the candidate factor that can suppress the reverberation to the greatest extent while minimizing the distortion of the target sound source signal, i.e., the factor that obtains the optimal objective evaluation value, is determined and fixed as the predetermined gain suppression factor for the specific environment, which is used for subsequent online operation. The base frequency estimates of the last 20 frames are then combined into an observation sequence, and a least square linear regression analysis is performed on this sequence to obtain the slope of the base frequency over time And 5 such observation sequences are continuously calculated, if the absolute values of the slopes of all 5 sequences exceed a slope threshold , and their signs remain consistent, the system will determine that the sound source is in motion, wherein the slope threshold is determined by the maximum radial motion speed of the sound source in the target scene , which is calculated by the Doppler effect formula , wherein is the typical base frequency of the sound source, is the speed of sound; once the sound source is determined to be in motion, the system will stop the cross-correlation operation and switch to energy envelope tracking, which is continuously generated by calculating the short-time root mean square energy of the subsequent audio stream with a frame length of 20 ms and a frame shift of 5 ms, and this energy envelope sequence is then used for time difference of arrival calculation of the dual-microphone array.

[0055] In addition, the application of deep neural networks in background noise recognition and the joint decision mechanism, the procedure is as follows, each captured source waveform segment is input into the cross-correlation operation at the same time, also input into an offline trained one-dimensional convolutional neural network, the network performs forward inference on the input segment, and outputs a posterior probability vector , wherein is the total number of predefined background noise categories; the only condition for an acoustic event to be finally determined as background noise is that the disorder index calculated based on the reverberation feature sequence is higher than its predefined disorder threshold , and the maximum value in the posterior probability vector output by the aforementioned neural network is also higher than a predetermined confidence threshold ; the predetermined confidence threshold here is determined by selecting the confidence point corresponding to the preset false positive rate of one in ten thousand on the receiver operating characteristic curve drawn on the validation set during the offline verification phase of the network, thereby fixing the decision preference of the classifier in the procedure.

[0056] It is obvious to those skilled in the art that the present application is not limited to the details of the above exemplary embodiments, and can be implemented in other specific forms without departing from the spirit or essential characteristics of the present application.

[0057] Finally, it should be noted that the above examples are merely intended to illustrate the technical solutions of the present application and not to limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. A multi-source soundfield localization and separation algorithm, characterized in that, The algorithm comprises the following steps: Step a, acquiring an audio stream and continuously monitoring the energy of the audio stream; when the energy of the audio stream first exceeds a dynamic energy threshold dynamically adjusted according to a real-time average energy of an environmental background noise, an initial audio segment of a predetermined length from the time point is intercepted and determined as a source-born waveform segment, the source-born waveform segment being an initial waveform characteristic of the sound source itself; Step b, using the source-born waveform segment as a matching template, performing a sliding cross-correlation operation on the audio stream after the source-born waveform segment to detect one or more distortion-related peaks that are highly related to the waveform characteristics of the source-born waveform segment and have energy attenuation in the time domain, the time delay and amplitude attenuation of the distortion-related peaks being path and environmental information of sound propagation; Step c, determining the source-born waveform segment and the audio signal immediately following the source-born waveform segment as direct sound, and identifying the audio signal segment corresponding to the one or more distortion-related peaks as reverberation of the direct sound; performing sound source positioning processing on the audio stream according to the identification result of the direct sound and the reverberation in the time domain; when an energy mutation that exceeds the dynamic energy threshold and has a cross-correlation lower than a predetermined cross-correlation threshold with the source-born waveform segment currently being used is detected in the audio stream, a new source-born waveform segment is generated, and the steps of detecting distortion-related peaks, determining direct sound and reverberation, and performing sound source positioning processing are performed in parallel on the new source-born waveform segment.

2. The multi-source sound field localization and separation algorithm of claim 1, wherein, After the step of detecting distortion-related peaks and before the step of sound source positioning processing, the following steps are further included: continuously tracking the fundamental frequency of the source-born waveform segment and the audio signal immediately following the source-born waveform segment to obtain a trajectory of the fundamental frequency changing over time; determining whether the fundamental frequency trajectory presents a continuous monotonic change; if not, the step of sound source positioning processing is continued; if yes, the step of sliding cross-correlation operation is suspended, and a cooperative tracking method based on energy envelope and time difference of arrival is switched to, wherein a continuous energy envelope of the sound source is continuously tracked, and a time difference of arrival of the energy envelope is calculated using a double-microphone array to obtain the direction of the moving sound source.

3. The multi-source sound field localization and separation algorithm of claim 1, wherein, After the step of detecting distortion-related peaks and before the step of sound source positioning processing, the following steps are further included: based on the time position and / or amplitude of the one or more distortion-related peaks in the time domain, a reverberation feature sequence is constructed; an index of disorder degree of the reverberation feature sequence is calculated, the index of disorder degree being determined according to the variance of the time interval sequence and the monotonicity degree of the amplitude attenuation sequence in the reverberation feature sequence; when the index of disorder degree is lower than a predetermined disorder degree threshold, the step of sound source positioning processing is performed; when the index of disorder degree is higher than the predetermined disorder degree threshold, subsequent processing steps are suspended, and the current source-born waveform segment is identified as background noise.

4. The multi-source sound field localization and separation algorithm of claim 3, wherein, After the index of disorder degree is higher than the predetermined disorder degree threshold and the current source-born waveform segment is identified as background noise, the current source-born waveform segment is marked as a background noise template, and in subsequent audio stream monitoring, when the energy exceeds the dynamic energy threshold, the background noise template is preferentially used for matching with the energy mutation, and if the matching is successful, subsequent processing is not performed.

5. The multi-source sound field localization and separation algorithm of claim 1, wherein, The predetermined length of the initial audio segment is five to ten milliseconds.

6. The multi-source sound field localization and separation algorithm of claim 1, wherein, The sliding cross-correlation operation is accelerated by fast Fourier transform.

7. The multi-source sound field localization and separation algorithm of claim 1, wherein, The sound source localization process comprises: when a double microphone array is used, calculating the azimuth of the reflecting surface from the time delay difference of the first detected distortion dependent peak arriving at the two microphones of the double microphone array , calculating the azimuth of the reflecting surface The azimuth is calculated according to the following formula: wherein is the sound velocity, is the distance between the two microphones.

8. The multi-source sound field localization and separation algorithm of claim 1, wherein, The sound source separation processing includes applying a predetermined gain suppression factor for gain suppression to an audio signal segment identified as reverberation.

9. The multi-source sound field localization and separation algorithm of claim 1, wherein, The sound source separation processing includes subtracting energy of the reverberation from the mixed audio signal.

Citation Information

Patent Citations

  • Hybrid speech processing method, electronic equipment and computer readable medium

    CN120236599A

  • Position detection system, transmission device, reception device, position detection method and position detection program

    US20110116345A1