Neuromorphic auditory event based dereverberation processing method and system for security scenario abnormal voiceprint recognition

By introducing a channel-level dynamic refractory period mask window and adaptive leakage adjustment of membrane potential into the neuromorphic auditory system, the problem of performance degradation caused by delayed redundant pulses in high reverberation environments is solved, and abnormal voiceprint recognition with high accuracy and low power consumption is achieved.

CN122157695APending Publication Date: 2026-06-05YUESHENGDA (TIANJIN) INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
YUESHENGDA (TIANJIN) INTELLIGENT TECHNOLOGY CO LTD
Filing Date
2026-03-09
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing neuromorphic auditory technologies suffer from reduced recognition performance and increased false alarms due to delayed redundant pulses generated by multipath echoes in highly reverberant enclosed spaces. Current methods cannot effectively suppress these delayed redundant pulses.

Method used

In a purely asynchronous auditory discrete event sequence, the instantaneous pulse excitation density rate of the frequency channel is calculated, the direct real pulse cluster is identified and the frequency envelope baseline vector is generated. The channel-level dynamic refractory period mask window is set in combination with the physical characteristics of the deployment scenario to intercept delayed redundant pulses, and the pulses are classified by a deep cascaded pulse neural network.

Benefits of technology

It effectively filters out redundant pulses generated by physical echoes, improving the accuracy and robustness of abnormal voiceprint recognition, preventing false alarms, and maintaining the advantage of low power consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122157695A_ABST
    Figure CN122157695A_ABST
Patent Text Reader

Abstract

The application discloses a neuromorphic auditory event-based anti-reverberation processing method and system for security scene abnormal voiceprint recognition, which comprises the following steps: acquiring an asynchronous auditory event sequence in a high-reverberation space; anchoring an initial pulse cluster triggered by a direct sound wave and extracting the same as a reference feature; constructing a channel-level dynamic refractory period mask window based on the reference feature and an estimated reverberation duration; determining a newly-arrived pulse cluster similar to the reference feature as a delayed ghost pulse generated by an echo and intercepting and filtering the same within the mask window action period, so as to generate a pure pulse stream; and inputting the pure pulse stream into a pulse neural network for recognition. The application effectively filters out redundant pulses generated by physical echoes from a data source in a native asynchronous pulse domain, avoids high-cost waveform reconstruction, can effectively prevent the subsequent pulse neural network from making a misjudgment and power saturation due to redundant data, improves the accuracy, robustness and real-time performance of the system in the abnormal voiceprint recognition under a high-reverberation environment, and meanwhile, retains the low-power consumption advantage of event-driven sensing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to artificial intelligence and signal processing technology based on neuromorphic auditory events, and in particular to an anti-reverberation processing method and system for security scenarios. Background Technology

[0002] In intelligent security monitoring technology, auditory event detection and voiceprint recognition are effective supplements to visual monitoring, playing a crucial role, especially in the early warning and perception of sudden abnormal events such as glass breakage, cries for help, vehicle collisions, or shootings. Neuromorphic auditory event-driven sensors, such as cochlear implants, are becoming a cutting-edge technology in security acoustics applications due to their bio-inspired mechanisms. These sensors employ an asynchronous event-driven detection paradigm, not relying on periodic sampling of continuous audio waveforms. Instead, they use a cascaded analog bandpass filter bank simulating the biological basilar membrane to encode the dynamic energy changes of acoustic signals in the environment into discrete asynchronous pulse sequences in real time. When the environment is quiet or there are no significant acoustic changes, the sensor generates very little data; however, when a sudden abnormal sound occurs, the sensor can capture the edge dynamics of the sound signal with microsecond-level precision, exciting a high-temporal-resolution pulse stream. This paradigm has the advantage of extremely low power consumption, making it suitable for continuous, 24 / 7 deployment at the edge nodes of security networks.

[0003] However, in certain security scenarios, the acoustic characteristics of the physical space can severely interfere with the encoding process of discrete pulses. When such event-driven auditory systems are deployed in spaces with high reverberation and enclosed characteristics, such as underground parking garages, narrow tunnels, elevator shafts, large vacant warehouses, or indoor corridors with hard walls, the surrounding concrete walls, load-bearing columns, and hardened floors form high-intensity acoustic reflective surfaces. In these environments, when a real abnormal sound event (such as the sound of breaking glass) occurs, its direct sound wave first reaches the sensor, triggering a cluster of "initial pulse sequences" characterizing the event. Subsequently, the sound is reflected and delayed multiple times by the surrounding walls, floor, and ceiling, and its long-tailed echo continues to reach the sensor's acquisition end.

[0004] Neuromorphic sensors are sensitive to high-frequency dynamic energy changes but lack error correction mechanisms based on global continuous waveform integration. Therefore, when faced with these physical multipath echoes, the underlying hardware indiscriminately encodes a large number of "delayed redundant pulses" on the time axis. These pulses are similar in frequency spatial distribution but exhibit microscopic phase delays and amplitude attenuation. These delayed redundant pulses are generated together with the pulses representing the real event, forming a redundant pulse stream, which is then input into the subsequent Spiking Neural Network (SNN) used for classification.

[0005] Existing audio dereverberation models typically rely on time-frequency domain transformations and matrix operations on complete, continuous audio waveforms, such as short-time Fourier transforms or Mel-frequency spectrum calculations. This type of continuous-frame processing architecture, based on synchronous dense sampling, is incompatible with the asynchronous, discrete, event-driven paradigm of neuromorphic acoustic systems and cannot be directly applied to purely event-driven pulse domain computation chains. On the other hand, existing spiking neural network models, such as the Leakage Integral Discharge (LIF) neuron model, are designed to preserve short-term memory and enhance the feature correlation between preceding and following pulses due to their inherent mechanisms (such as membrane potential leakage attenuation and short-term synaptic plasticity). This results in the network not only failing to suppress echo noise with highly similar features but also treating delayed redundant pulses as continuous high-frequency homogeneous inputs, performing cross-channel spatiotemporal integration. As a result, in confined security spaces with strong reverberation, spiking neural networks can misidentify a single physical event as a dense, high-frequency burst of events within a short period, triggering numerous false alarms. This leads to network synaptic potential saturation, a sharp increase in computing power and power consumption, potentially causing data overload in the alarm channel, and ultimately reducing the accuracy and robustness of abnormal soundprint recognition. Therefore, designing an effective method for suppressing and filtering delayed redundant pulses within a framework that relies purely on asynchronous discrete pulse data without waveform reconstruction remains a technical problem to be solved in this field. Summary of the Invention

[0006] One aspect of the present invention is to provide an anti-reverberation processing method and system based on neuromorphic auditory events for abnormal voiceprint recognition in security scenarios, so as to solve the technical problem that the recognition performance of existing neuromorphic auditory technology is reduced due to the delay redundant pulses generated by multipath echoes in a strongly reverberant confined space.

[0007] To achieve the above objectives, the first aspect of the present invention provides an anti-reverberation processing method based on neuromorphic auditory events for abnormal voiceprint recognition in security scenarios, comprising the following steps:

[0008] Step 1: Obtain the pure asynchronous auditory discrete event sequence, and parse each discrete event in the pure asynchronous auditory discrete event sequence into a digital tuple data packet containing a timestamp, polarity information, and frequency channel address;

[0009] Step 2: Calculate the instantaneous pulse excitation density rate of the frequency channel address. When the instantaneous pulse excitation density rate exceeds the preset density threshold, it is determined that the direct real original pulse cluster induced by the direct original sound wave has been captured. Set the arrival time of the energy peak of the direct real original pulse cluster as the global reference timestamp of the first wave occurrence, extract its spatial distribution in the frequency channel address dimension to determine the active channel set, and generate the direct wave original frequency envelope baseline vector.

[0010] Step 3: Based on the physical characteristics of the deployment scenario, determine the estimated reverberation duration; taking the global reference timestamp of the first wave occurrence as the starting point, apply a channel-level dynamic refractory period mask window with an effective duration determined based on the estimated reverberation duration to the activated channel set; measure the similarity between the newly arriving pulse clusters within the effective period of the channel-level dynamic refractory period mask window and the original frequency envelope baseline vector of the direct wave; when the measurement result meets the preset similarity condition, determine that the newly arriving pulse clusters are delayed redundant pulses and intercept them to generate a denoised discrete pulse data stream;

[0011] Step 4: Convert the denoised discrete pulse data stream into a two-dimensional topological time feature map;

[0012] Step 5: Input the two-dimensional topological time feature map into a deep cascaded spiking neural network for classification, and generate an early warning signal based on the classification result.

[0013] Further, in step two, the instantaneous pulse excitation density rate of the frequency channel address is calculated by setting a timing sliding monitoring window, and the length of the timing sliding monitoring window ranges from 1 microsecond to 5000 microseconds; the instantaneous pulse excitation density rate exceeding the preset density threshold means that the instantaneous pulse excitation density rate of one frequency channel address or multiple frequency-adjacent associated frequency channel addresses exceeds the preset density threshold within the time length of the timing sliding monitoring window.

[0014] Preferably, the channel-level dynamic refractory period mask window adopts an asymmetric mask adaptive shrinkage strategy based on the cross-band acoustic energy dissipation gradient; specifically, according to the frequency band height of each frequency channel address contained in the original frequency envelope baseline vector of the direct wave, a set of frequency-related mask duration parameters are generated, and according to the parameters, a channel-level dynamic refractory period mask window with an action period shorter than the estimated reverberation time is set for frequency channel addresses with frequencies higher than a preset frequency threshold, and a channel-level dynamic refractory period mask window with an action period equal to the estimated reverberation time is set for frequency channel addresses with frequencies lower than the preset frequency threshold.

[0015] Optionally, in step three, the similarity measurement is calculated using cosine similarity; the measurement result satisfies the preset similarity condition, specifically meaning that the result of the similarity measurement is greater than or equal to the preset similarity threshold. At this time, the newly arrived pulse cluster is determined to be a delayed redundant pulse and interception and discarding are performed.

[0016] Furthermore, the method also includes a step of implementing adaptive decay suppression of network membrane potential between step four and step five. This step includes: calculating the number of pulses in the denoised discrete pulse data stream whose excitation intensity is lower than a set background noise energy threshold to assess the residual background noise level; and adaptively adjusting the membrane potential leakage integral constant of neurons in the deep cascaded spiking neural network according to the background noise level, and increasing the basic leakage rate of neuron membrane potential when the background noise level is detected to be greater than a preset background noise threshold.

[0017] Optionally, in step four, the denoised discrete pulse data stream is converted into the two-dimensional topological time feature map by an exponential smooth decay algorithm based on the pulse arrival time.

[0018] Furthermore, in step three, the estimated reverberation duration is set as follows: it is calculated using the formula T_reverb = (2d / v) * k, where d is the distance between the sensor and the nearest acoustic reflector in the scene, v is the speed of sound propagation in the spatial medium, k is the environmental diffraction redundancy ratio coefficient, and T_reverb is the estimated reverberation duration.

[0019] Optionally, the deep cascaded spiking neural network is a spiking convolutional neural architecture based on leakage integral and firing mechanism, or a spiking neural network architecture with a time-dependent synaptic feature self-attention mechanism.

[0020] Furthermore, in step five, the early warning signal includes the classification result and the global reference timestamp of the first wave occurrence.

[0021] A second aspect of the present invention provides an anti-reverberation processing system based on neuromorphic auditory events for abnormal voiceprint recognition in security scenarios, comprising:

[0022] The purely asynchronous auditory acquisition module is used to acquire a sequence of purely asynchronous auditory discrete events and parse each discrete event in the sequence of purely asynchronous auditory discrete events into a digital tuple data packet containing a timestamp, polarity information and frequency channel address;

[0023] The baseline feature extraction module is used to calculate the instantaneous pulse excitation density rate of the frequency channel address. When the instantaneous pulse excitation density rate exceeds a preset density threshold, it is determined that the direct real original pulse cluster induced by the direct original sound wave has been captured. The arrival time of the energy peak of the direct real original pulse cluster is set as the global baseline timestamp of the first wave occurrence. Its spatial distribution in the frequency channel address dimension is extracted to determine the active channel set, and the baseline vector of the direct wave original frequency envelope is generated.

[0024] The mask-side suppression and filtering module is used to determine the estimated reverberation duration based on the physical characteristics of the deployment scenario; starting from the global reference timestamp of the first wave occurrence, it applies a channel-level dynamic refractory period mask window with an effective duration determined based on the estimated reverberation duration to the activated channel set; it performs a similarity measurement between the newly arriving pulse clusters within the effective period of the channel-level dynamic refractory period mask window and the original frequency envelope baseline vector of the direct wave; when the measurement result meets the preset similarity condition, it determines that the newly arriving pulse clusters are delayed redundant pulses and performs interception to generate a denoised discrete pulse data stream;

[0025] An adaptive two-dimensional topology and noise reduction reconstruction module is used to convert the denoised discrete pulse data stream into a two-dimensional topological time feature map;

[0026] The cascaded spiking neural network classification and early warning module is used to input the two-dimensional topological time feature map into a deep cascaded spiking neural network for classification, and generate an early warning signal based on the classification result.

[0027] The beneficial effects of this invention are as follows: This invention implements a channel-level dynamic refractory period mask window within the native asynchronous discrete pulse data domain, avoiding the computationally expensive waveform reconstruction process. This provides an effective technical means for neuromorphic auditory systems to identify abnormal sounds in extreme security environments with high reverberation and strong reflection, such as underground parking garages. This method can effectively filter out redundant pulses generated by physical echoes from the data source, thereby effectively preventing subsequent pulse neural networks from making misjudgments and saturating computational power. This improves the accuracy, robustness, and real-time performance of abnormal voiceprint recognition, while retaining the low-power advantage of event-driven sensors, effectively preventing the continuous occurrence of false alarms. Attached Figure Description

[0028] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0029] Figure 1 This is a flowchart of an anti-reverberation processing method based on neuromorphic auditory events for abnormal voiceprint recognition in security scenarios, provided by an embodiment of the present invention.

[0030] Figure 2 This is a schematic diagram illustrating the mechanism of delay redundant pulse generation under strong reverberation conditions.

[0031] Figure 3 This is a schematic diagram illustrating the working principle of the channel-level dynamic refractory period mask window in the embodiments of the present invention.

[0032] Figure 4 This is a structural block diagram of an anti-reverberation processing system based on neuromorphic auditory events for abnormal voiceprint recognition in security scenarios, provided by an embodiment of the present invention.

[0033] Figure 5 This is a visualization of the pulse cluster frequency envelope similarity determination process of the present invention.

[0034] Figure 6 This is a frequency domain illustration of the cross-band asymmetric mask adaptive shrinkage strategy of the present invention.

[0035] Figure 7 This is a schematic diagram of the adaptive leakage regulation mechanism of LIF neuron membrane potential in this invention.

[0036] Figure 8 This is a diagram illustrating the sensor deployment and sound wave propagation path in an underground parking garage security scenario according to the present invention.

[0037] Figure 9 This is a comparison chart showing the accuracy of abnormal voiceprint recognition under different reverberation intensities.

[0038] Figure 10 This is a comprehensive comparison chart of the present invention and other solutions in terms of system power consumption and processing latency. Detailed Implementation

[0039] To make the objectives, technical solutions, and advantages of the present invention clearer, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0040] Example 1

[0041] Please see Figure 1 This embodiment provides an anti-reverberation processing method based on neuromorphic auditory events for abnormal voiceprint recognition in security scenarios. This method can be applied to intelligent monitoring equipment deployed in enclosed security spaces with high reverberation characteristics. The method aims to suppress redundant pulse data generated by acoustic multipath reflections at the data acquisition front end, thereby improving the accuracy of subsequent abnormal voiceprint recognition by the spiking neural network.

[0042] Specifically, refer to Figure 1 The process shown includes the following steps:

[0043] Step S1: Obtain a sequence of purely asynchronous auditory discrete events in a high-reverberation security space and perform multi-frequency channel mapping analysis.

[0044] This step is performed in a system configured for a specific security scenario characterized by high acoustic reflection and reverberation, such as an underground parking garage with concrete walls and floors, a narrow tunnel, an elevator shaft with metal interior walls, or a large vacant warehouse. The system utilizes a neuromorphic auditory event-driven sensor to continuously monitor ambient acoustic signals. The sensor's internal structure mimics the basilar membrane of the cochlea, containing a multi-stage cascaded active analog bandpass filter bank network, such as a Gammatone filter bank with 64 independent channels. When the energy change of the ambient sound signal, i.e., the slope of the sound pressure change, exceeds the sensor's preset hardware neuromorphic trigger threshold, the sensor's internal logarithmic compression circuit works in conjunction with the filter bank to encode this dynamic change in real time into a one-dimensional initial action potential pulse sequence. This process is asynchronous, generating data only when there is a significant change in sound, thus achieving extremely low standby power consumption.

[0045] like Figure 8 As shown, in the underground parking garage security scenario, concrete walls 801 form a closed space, and several load-bearing columns 802 are installed within the scenario for structural support. Neuromorphic auditory sensors 803 are installed on the surface of each load-bearing column 802, forming a sensor network covering the entire parking garage space. Vehicles 805 are parked within the scenario, and the hardened ground 806 forms a sound wave reflecting surface. When an abnormal sound event 804 occurs in the parking garage, the sound waves propagate to the sensor 803 along multiple paths: the direct sound wave propagates along the shortest path from the sound source 804 to the sensor 803; the first reflected echo propagates from the sound source 804 to the concrete wall 801, undergoes a first reflection, and then reaches the sensor 803; the second reflected echo is reflected sequentially from the sound source 804 onto two concrete walls 801 before reaching the sensor 803. The distance d between the sensor 803 and the nearest wall 801 is 25 meters. Due to the high reflectivity of the concrete walls 801 and the hardened ground 806, the underground garage constitutes a typical high-reverberation enclosed space. The direct sound waves and multiple reflected echoes superimpose in time, forming a complex reverberation interference environment.

[0046] Furthermore, the system parses each discrete event in the initial action potential pulse sequence. This parsing process follows the Address Event Representation (AER) protocol, converting each pulse event into a structured digital tuple data packet in the format [t, p, ch]. Here, t is the absolute timestamp of the pulse triggering, with an accuracy down to the microsecond level, recording the precise moment the event occurred; p is polarity information, used to characterize whether the energy envelope increases or decreases instantaneously, for example, +1 can represent an energy increase, and -1 represents an energy decrease; ch is the frequency channel address, an integer that uniquely identifies which specific frequency band in the filter bank network triggered the pulse, thus preserving the spectral information of the sound signal.

[0047] Step S2 involves performing spatiotemporal clustering on the real pulse clusters induced by the direct original sound wave and extracting high-dimensional anchoring features of the joint frequency domain benchmark.

[0048] See Figure 2 The diagram illustrates that after a real abnormal sound event 101 occurs, its direct sound wave 102 first reaches the sensor, generating an initial direct real original pulse cluster 201. Subsequently, the echo wave 104 generated by reflective surfaces such as the wall 103 reaches the sensor and is incorrectly encoded as a delayed redundant pulse 202. The goal of this step is to accurately identify the direct real original pulse cluster 201 in the pure pulse domain without relying on traditional time-frequency transformation.

[0049] Specifically, the system processes received AER data packets asynchronously. The system sets a very short initial timing sliding monitoring window, the length of which can be configured according to the response speed requirements of the application scenario; preferably, it can be set between 1 microsecond and 5000 microseconds. Within this monitoring window, the system dynamically and continuously calculates the instantaneous pulse excitation density rate (PER), i.e., the pulse count value per unit time, within all frequency channel addresses (ch). When the calculated instantaneous pulse excitation density rate in a single frequency channel, or multiple adjacent frequency channels (e.g., 3 to 5 consecutive channels in the frequency address encoding), rises sharply within a very short time and exceeds a preset density threshold, the system determines that it has successfully captured a "direct-reaching real original pulse cluster" directly propagated from the physical sound source. This threshold can be a fixed value based on historical background noise statistics or a dynamic value adaptively adjusted according to recent environmental noise levels.

[0050] After capturing the pulse cluster, the system marks this batch of pulses that are continuous in time and discharge at high density as a whole. For the consistency of subsequent processing, the system calculates the microsecond time center axis when the energy peak of the pulse cluster arrives and sets it as the "global reference timestamp of the first wave occurrence", denoted as T_0. Finally, the system extracts the spatial distribution characteristics of the direct-to-real original pulse cluster in the frequency address dimension, forming an intensity spectrum, and combines it with its polarity response combination to define it as a high-dimensional "activation channel set". Further, several core channels with the highest excitation rate are selected from this set, and their excitation intensities are weighted and combined or directly spliced ​​to obtain a "direct-to-wave original frequency envelope baseline vector". This vector mathematically represents the core spectral characteristics of the original sound event and serves as a digital reference and template for identifying and filtering delayed ghost pulses in subsequent steps.

[0051] Step S3: Based on the channel-level dynamic refractory period mask window, the delayed redundant pulses are suppressed and purified.

[0052] This step is the core of the invention, aiming to actively intercept redundant pulses generated by echoes by utilizing the baseline features extracted in the previous step. See also Figure 3 The diagram illustrates the working principle of this step.

[0053] First, the system needs to determine a reasonable shielding period. This period is estimated based on the physical acoustic characteristics of the deployment scenario. Specifically, the system sets an "estimated reverberation duration," denoted as T_reverb. This duration can be directly input using prior knowledge; for example, for a warehouse of known size, an empirical value can be preset. Alternatively, this constant can be obtained through adaptive calculation. The formula for calculating the estimated reverberation duration is T_reverb = (2d / v) * k. In this formula, d represents the straight-line distance between the sensor location and the nearest major acoustic reflector (such as a wall) in the scene, v represents the average propagation speed of sound in the spatial medium (usually air), approximately 343 m / s, and k is the environmental diffraction redundancy scaling factor, for example, between 1.1 and 1.2 times the base value, to ensure coverage of the vast majority of significant echo paths. In typical underground parking garages or tunnel environments, the T_reverb value can be set between thirty milliseconds and five hundred milliseconds.

[0054] In specific implementation, such as Figure 3 As shown, starting from the "global reference timestamp of the first wave occurrence" T_0 established in step S2 above, the system applies a time-effective "channel-level dynamic refractory period mask window" 301 to the processing nodes within the "active channel set" and its harmonic associated channels. The effective duration of this mask window is T_reverb. During the effective period of this window (i.e., from T_0 to T_0 + T_reverb), when the sensor continuously outputs new pulse data due to receiving echo waves, the mask-side suppression filtering module captures these newly arriving pulse clusters and performs a rapid feature scan. The system uses a fast judgment method, such as calculating the cosine similarity between the frequency envelope vector of the new pulse cluster to be judged and the cached "direct wave original frequency envelope baseline vector", or calculating the overlap between the two in the characteristic frequency band, to perform multi-dimensional dynamic similarity measurement.

[0055] When the similarity measurement result is greater than or equal to a preset "similarity threshold" (e.g., it can be set empirically in the range of 0.8 to 0.95, and is set to 0.85 in this embodiment), and the timestamps of the newly arrived pulse clusters fall completely within the effective time range of the channel-level dynamic refractory period mask window, the system can determine that the newly arrived pulse clusters with highly similar spectral structures are "delayed redundant pulses" 202 caused by spatial multipath reflection. Once the determination is made, before any data stream is sent to the subsequent spiking neural network, the control command will perform forced interception, discarding from memory, and setting the excitation intensity to zero in the data stream for the data queue marked as delayed redundant pulses. Through this mechanism, the final generated data stream is a denoised discrete pulse data stream 302 containing only the edge information of the initial burst event.

[0056] like Figure 5 The figure illustrates the complete process of determining the cosine similarity of the frequency envelope of a pulse cluster. The upper part of the figure shows the frequency envelope baseline vector extracted after the direct-arrival pulse cluster is anchored by the instantaneous excitation density rate exceeding the threshold. This vector covers 64 frequency channels (ch=1 to ch=64), and the excitation intensity of each channel is represented by a grayscale bar chart, reflecting the energy distribution characteristics of the direct sound in each frequency channel. The lower part of the figure shows the frequency envelope vector of a newly arrived pulse cluster to be judged within the mask window's validity period, located in the same 64 channels. The two sets of vectors are connected by dashed lines to visually demonstrate the channel-by-channel excitation intensity comparison. After cosine similarity calculation, the similarity between the two sets of vectors, cos(T) = 0.94, is higher than the preset judgment threshold of 0.85. Therefore, this pulse cluster is judged as a delayed redundant pulse generated by physical echo, and an interception operation is performed. This judgment mechanism can accurately identify and filter out redundant pulses with high frequency structural similarity to the direct wave while retaining new event pulses from different sources.

[0057] Preferably, to achieve refined and differentiated processing of echoes at different frequencies, the channel-level dynamic refractory period mask window adopts a variable duration. Considering that the energy attenuation rate of high-frequency sound waves in air and wall reflections is much faster than that of low-frequency sound waves, this invention also provides an asymmetric mask adaptive contraction strategy based on the cross-band sound wave energy dissipation gradient. Specifically, after capturing the "direct wave original frequency envelope baseline vector," the system calculates and generates a set of "frequency-related mask duration parameters" based on the frequency band corresponding to each frequency channel address ch in the vector. According to these parameters, a rapidly decaying dynamic refractory period mask window is set for address channels in the high-frequency band, and its duration can be much smaller than the average T_reverb, for example, only 30% of it. The advantage of doing so is that it can quickly restore the ability to perceive new, real high-frequency events. At the same time, a mask window with a longer duration is set for address channels in the low-frequency band, and its duration can be extended to the upper limit of T_reverb, thereby effectively suppressing long-lasting low-frequency standing wave echoes that are easily formed in confined spaces. Accordingly, when performing similarity measurement, the judgment weights are also dynamically adjusted over time: for pulse clusters to be evaluated with a long delay time, the coherence comparison weight between them and the baseline vector in the low-frequency components will be adaptively increased.

[0058] like Figure 6 The figure illustrates the asymmetric shrinkage strategy of the mask window in the frequency-time two-dimensional space. The horizontal axis represents the time range from the direct wave anchoring time T0 to T0+T_reverb, and the vertical axis represents the 64 frequency channels (ch=1 for low-frequency channels and ch=64 for high-frequency channels). The effective duration of the mask window for each channel is represented by rectangular areas of different gray levels: low-frequency channels (ch1 to ch19, corresponding to frequencies below 500Hz) are marked in dark gray, with a mask duration covering the complete reverberation time, i.e., 100% T_reverb, achieving continuous suppression of the low-frequency band; mid-frequency channels (ch20 to ch48) are marked in medium gray, with the mask duration linearly decreasing from 100% to 30% T_reverb; high-frequency channels (ch49 to ch64, corresponding to frequencies above 4kHz) are marked in light gray, with a mask duration of only 30% T_reverb, achieving rapid recovery of the high-frequency band. This asymmetric shrinkage design conforms to the physical law that low-frequency reverberation decays slowly and high-frequency reverberation decays quickly in indoor acoustics. While effectively filtering out redundant pulses in each frequency band, it maximizes the retention of the time resolution capability of the high-frequency channel.

[0059] Step S4: The purified anti-reverberation pulse current is represented by a two-dimensional feature space evolution and network membrane potential adaptive decay suppression is implemented.

[0060] First, to enable the discrete pulse stream to be accepted by the subsequent image-processing neural network, the system needs to convert it into a spatiotemporal topological feature representation. Specifically, an exponential smooth decay algorithm based on the pulse arrival time can be used. Whenever the system receives a real pulse event [t, p, ch] after filtering in step S3, an active memory parameter containing its recent local temporal history is generated for the event using an exponential time decay kernel function. For example, in a two-dimensional coordinate system, with the frequency channel ch as one axis and time as the other, a peak is generated at the position (ch, t), which decays exponentially with time. By fusing the decay surfaces generated by all pulses, a composite two-dimensional topological temporal feature map is generated. This feature map condenses the spatiotemporal dynamic information of the pulse stream and serves as the denoised base tensor input to the subsequent spiking neural network.

[0061] Further, optionally, this invention provides a model response filtering strategy based on the residual background noise intensity of the reverberant environment. This mechanism dynamically assesses the level of residual background noise not completely filtered out by the mask window by continuously calculating the number of weak diffuse pulses with excitation intensity below a set background noise energy threshold extracted from the denoised discrete pulse sequence. Based on this assessment result, the membrane potential leakage integral constant of all neurons in the downstream spiking neural network is adaptively adjusted. For example, in the LIF neuron model, the differential equation of the membrane potential V is dV / dt = -(V - V_rest) / tau + I_syn. Here, t is time, V is the neuron's membrane potential, V_rest is the resting potential, I_syn is the synaptic input current, and tau is the membrane time constant, which determines the leakage rate. When a relatively high background noise level is detected due to diffuse reflection, the system autonomously decreases the tau value, i.e., increases the baseline leakage rate of the neuron's membrane potential. This allows the non-coherent, weak residual delay redundant pulses that are not intercepted by the mask window to accumulate on the membrane potential and dissipate due to faster capacitor discharge, making it more difficult to reach the threshold for triggering pulse emission. This further enhances the system's ability to resist false triggering.

[0062] like Figure 7The figure illustrates the evolution of the membrane potential V(t) of a LIF neuron over time under different leakage time constants (tau). The triangles at the bottom of the figure represent redundant pulse inputs remaining after being intercepted by the mask window. The solid curve corresponds to the case with a normal tau value (tau=5.0). Due to the slow leakage rate, the increase in membrane potential from each pulse input does not sufficiently decay before the next input arrives, causing the membrane potential to gradually accumulate and eventually exceed the firing threshold V_th, resulting in false triggering (marked with an cross in the figure). The dashed curve corresponds to the case after adaptively reducing tau (tau=1.0). Due to the significantly faster leakage rate, the membrane potential rapidly decays to a lower level after each pulse input. Even with multiple consecutive residual pulse inputs, the membrane potential cannot accumulate to the firing threshold, thus effectively suppressing residual redundant pulses. The core equation of this adaptive leakage regulation mechanism is dV / dt=-(V-V_rest) / tau+I_syn. By dynamically adjusting the tau parameter according to the residual noise level, the leakage rate is increased in high-noise environments, providing subsequent fault tolerance for the mask window.

[0063] Step S5: High-resolution inference and early warning decision-making for abnormal voiceprints are performed using a deep cascaded pulse neural network.

[0064] The two-dimensional pure topological temporal feature tensor stream, generated by the aforementioned steps and containing only initial physical event information, is serially input into a pre-trained deep cascaded spiking neural network model. This model can employ a spiking convolutional neural architecture based on the leaky integral and discharge (LIF) mechanism, or an SNN architecture with a time-dependent synaptic self-attention mechanism. Since the input data has already effectively filtered out redundant pulses caused by environmental reflections at the front end, the neural network no longer needs to bear a large computational load for denoising and dereverberation. Therefore, all computational resources can be used to efficiently perform spatiotemporal correlation analysis on the envelope edges of real acoustic events, such as identifying the high-frequency impact characteristics unique to glass breakage events or the low-frequency muffled characteristics of vehicle collision events.

[0065] In the network's classification output layer, each predefined type of abnormal event (such as broken glass, cries for help, or vehicle collision) corresponds to one or a group of output neurons. These output neurons execute a competition mechanism for membrane potential firing rates. When the cumulative membrane potential of a neuron representing a specific abnormal category crosses the preset firing threshold first due to receiving a pure pulse stream and continues to fire pulses, the system can establish a high-confidence classification result for that acoustic event. Simultaneously, the system can obtain the "first wave occurrence global reference timestamp" T_0 from the identified initial pulse cluster, obtaining a microsecond-level time stamp of the first wave center of the event. Finally, the system combines the classification result with the precise timestamp and sends an immediate warning signal to the security integration backend. This mechanism effectively prevents false repeating alarms caused by strong echo environments and avoids system computing power saturation or alarm channel information blockage caused by a large influx of ghost pulses.

[0066] Example 2

[0067] Please see Figure 4 This embodiment provides an anti-reverberation processing system based on neuromorphic auditory events for abnormal voiceprint recognition in security scenarios. This system is a hardware or hardware-software combination implementation of the aforementioned method. The system 400 includes:

[0068] The purely asynchronous auditory acquisition module 401 is used to acquire changes in ambient sound pressure using a neuromorphic auditory event-driven sensor in a closed security scenario with high echo reverberation, and to encode these changes in real time into discrete auditory pulse sequences following the AER format.

[0069] The reference feature extraction module 402 is connected to the purely asynchronous auditory acquisition module 401 and receives the pulse sequence output by it. Internally, it contains logic circuits or software algorithms for high-density burst screening of the acquired pulse sequence. It determines the "first wave occurrence global reference timestamp" T_0 of the initial sound wave by calculating the instantaneous pulse excitation density rate, and extracts the "direct wave original frequency envelope baseline vector" from this initial pulse cluster as the reference for subsequent comparison.

[0070] The mask-side suppression filtering module 403, connected to the reference feature extraction module 402, receives a global reference timestamp and a baseline vector. This module calculates the estimated reverberation duration T_reverb based on preset or real-time calculated acoustic field physical environment parameters, and applies a "channel-level dynamic refractory period mask window" with an effective duration of T_reverb to the active frequency channels contained in the baseline vector. This module compares newly arriving pulse clusters within the mask window with the stored "direct wave original frequency envelope baseline vector," identifying pulses with a similarity higher than a preset threshold as delayed redundant pulses, and performing interception and discarding processing, outputting a denoised discrete pulse data stream.

[0071] An adaptive two-dimensional topology and noise reduction reconstruction module 404, connected to the mask-side suppression filtering module 403, receives a clean pulse stream. Its function is to transform the filtered discrete pulse stream into a two-dimensional spatiotemporal topological feature map using spatiotemporal feature transformation methods such as exponential decay kernels. Optionally, this module also includes a residual noise evaluation unit that dynamically generates control signals based on the residual noise level to adaptively adjust the membrane potential leakage parameters of downstream network neurons.

[0072] The spiking neural network classification cascaded early warning module 405 is connected to the adaptive 2D topology and noise reduction reconstruction module 404 and receives a 2D topological feature map. Internally, it integrates a pre-trained deep spiking neural network model. This module is used to classify and identify the input feature map. When the output neuron representing a specific abnormal event category is activated, it generates a single abnormal event early warning signal with high confidence and accurate time labels, and transmits this signal to an external security management platform or alarm system.

[0073] In summary, this invention innovatively introduces a channel-level dynamic refractory period mask window and an optional membrane potential adaptive leakage adjustment mechanism within the native asynchronous discrete pulse data domain of the neuromorphic auditory system, avoiding the computationally expensive waveform reconstruction process of traditional solutions. This invention can accurately and efficiently filter out a large number of redundant pulses generated by physical echoes from the data source, thereby preventing subsequent spiking neural networks from misjudgments and computational saturation due to data contamination. It significantly improves the accuracy, robustness, and real-time performance of abnormal voiceprint recognition in extreme security environments with high reverberation and strong reflection, such as underground parking garages, while fully retaining the low-power core advantage of event-driven sensors, effectively preventing the continuous occurrence of false alarms.

[0074] like Figure 9As shown, under test conditions where the reverberation time gradually increased from 50ms to 500ms, a comparative experiment was conducted on the abnormal voiceprint recognition accuracy of three schemes: the method of this invention, the traditional waveform reconstruction de-reverberation combined with a spiking neural network, and the original spiking neural network without de-reverberation processing. The method of this invention achieved a recognition accuracy of 98.5% with a reverberation time of 50ms. When the reverberation time increased to 500ms, the accuracy only slowly decreased to 91.2%, with an overall decrease of approximately 7.3 percentage points, demonstrating excellent anti-reverberation robustness. In contrast, the traditional waveform reconstruction de-reverberation combined with a spiking neural network scheme, under the same conditions, saw its accuracy decrease from 95.0% to 72.3%, a decrease of 22.7 percentage points; while the accuracy of the original spiking neural network without de-reverberation processing dropped sharply from 90.2% to 45.8%, a decrease of as much as 44.4 percentage points. As can be seen from the area filled between the three curves in the figure, the performance advantage of the method of the present invention compared with the other two schemes becomes more and more significant as the reverberation intensity increases, which fully verifies that the anti-reverberation processing method based on neuromorphic auditory events of the present invention can still maintain a high recognition accuracy in high reverberation environments.

[0075] like Figure 10 As shown, a comprehensive comparison of the method of this invention, the traditional FFT déresonance scheme, the deep learning déresonance scheme, and the non-déresonance processing scheme is presented in terms of average power consumption and average processing latency. Regarding power consumption, the average power consumption of the method of this invention is only 12mW, slightly higher than the 8mW of the non-déresonance processing scheme, but far lower than the 85mW of the traditional FFT déresonance scheme and the 320mW of the deep learning déresonance scheme, reducing power consumption by approximately 85.9% and 96.3%, respectively. Regarding processing latency, the average processing latency of the method of this invention is 0.8ms, close to the 0.3ms level of the non-déresonance processing scheme, while the latency of the traditional FFT déresonance scheme is 15ms, and the latency of the deep learning déresonance scheme is as high as 45ms. In summary, the method of this invention introduces effective déresonance processing capabilities while generating only minimal additional power consumption and latency overhead, achieving an excellent balance between power consumption and real-time performance, and fully preserving the inherent low power consumption and low latency advantages of event-driven neuromorphic sensors.

[0076] Example 3

[0077] This embodiment further illustrates the implementation process and necessity of the method of the present invention through a specific security application scenario. In a large, multi-story underground parking garage environment, the space, composed of concrete walls, floors, and load-bearing columns, exhibits typical characteristics of strong acoustic reflection and long reverberation time. A security system based on neuromorphic auditory event-driven sensors is deployed in this scenario to monitor abnormal sounds such as vehicle collisions or tire blowouts. In this environment, when a real vehicle tire blowout occurs, the resulting short, high-energy sound is reflected multiple times in the space, causing the sensor to continuously receive echoes with gradually decaying energy but highly similar spectral characteristics for hundreds of milliseconds after receiving the direct sound wave. If the raw pulse stream generated by the sensor is directly fed into a subsequent spiking neural network for classification, the network will misclassify this series of pulses as multiple independent impact events occurring consecutively within a short period, thus sending a large number of repetitive false alarms to the security center, severely impacting the reliability of the system. Because conventional dereverberation methods in existing technologies require reconstructing sparse asynchronous pulse data into dense continuous audio waveforms for processing, this introduces huge computational overhead and latency, completely violating the low-power, event-driven design intent of neuromorphic sensors, and therefore is not suitable for such purely event-driven security systems.

[0078] The specific processing flow of the method provided by this invention is as follows: When a tire bursting sound occurs, the purely asynchronous auditory acquisition module 401, deployed on the load-bearing column of the parking lot, encodes the acoustic event into a pulse sequence in AER format. The baseline feature extraction module 402 sets the sliding monitoring window length to 1000 microseconds, accurately captures the initial high-density pulse cluster generated by the direct sound wave by monitoring the pulse excitation density rate, records the center point of its occurrence time as the global baseline timestamp T_0, and extracts its frequency channel distribution features as the "direct wave original frequency envelope baseline vector".

[0079] Subsequently, the mask-side suppression filtering module 403 is activated. Based on the physical parameters of the parking lot environment, such as the distance d between the sensor and the nearest load-bearing wall being 25 meters, the system calculates the basic estimated reverberation time based on the speed of sound (approximately 343 m / s), multiplies it by, for example, an environmental diffraction redundancy scaling factor k of 1.2, and finally sets T_reverb to 175 milliseconds. Starting from T_0, the module applies a "channel-level dynamic refractory period mask window" with a duration of 175 milliseconds to the active frequency channels contained in the baseline vector. In the time after T_0, when subsequent pulse clusters generated by the echo arrive, for example, a pulse cluster arriving at T_0+50 milliseconds, the module calculates the cosine similarity between its frequency envelope and the stored baseline vector. If the calculation result is 0.92, which is higher than the preset similarity threshold of 0.85, the pulse cluster is determined to be a delayed redundant pulse and is directly discarded. This process continues within the 175-millisecond window period, effectively filtering out all redundant data generated by the echo. Meanwhile, the module can execute an asymmetric mask shrinkage strategy, such as applying a short mask of 60 milliseconds to the high-frequency channels (e.g., above 4kHz) caused by a tire blowout to quickly restore sensitivity to new high-frequency events, while applying a full long mask of 175 milliseconds to the low-frequency channels (e.g., below 500Hz) with longer reverberation times.

[0080] After filtering, the pure pulse stream containing only the initial tire blowout event characteristics is sent to subsequent modules. The adaptive 2D topology and noise reduction reconstruction module 404 converts it into a 2D topological feature map and inputs it into the spiking neural network classification cascaded early warning module 405. Since the input data no longer contains reverberation interference, the network can accurately identify the event as "tire blowout" and generate a unique early warning signal by combining it with the timestamp T_0. Through the above processing, the security center ultimately receives only one accurate early warning about the tire blowout event, avoiding false alarms and thus improving the robustness of the recognition system in complex acoustic environments.

[0081] The above description is merely a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A neuromorphic auditory event based dereverberation processing method for security scenario anomaly voiceprint recognition, characterized in that, Includes the following steps: Step 1: Obtain the pure asynchronous auditory discrete event sequence, and parse each discrete event in the pure asynchronous auditory discrete event sequence into a digital tuple data packet containing a timestamp, polarity information, and frequency channel address; Step 2: Calculate the instantaneous pulse excitation density rate of the frequency channel address. When the instantaneous pulse excitation density rate exceeds the preset density threshold, it is determined that a direct real original pulse cluster induced by the direct original sound wave has been captured. The arrival time of the energy spike of the direct-to-real original pulse cluster is set as the global reference timestamp of the first wave occurrence. Its spatial distribution in the frequency channel address dimension is extracted to determine the active channel set, and the original frequency envelope baseline vector of the direct wave is generated. Step 3: Determine the estimated reverberation duration based on the physical characteristics of the deployment scenario; Starting from the global reference timestamp of the first wave occurrence, a channel-level dynamic refractory period mask window with an effective duration determined based on the estimated reverberation time is applied to the activated channel set. The similarity between the newly arriving pulse clusters within the effective period of the channel-level dynamic refractory period mask window and the original frequency envelope baseline vector of the direct wave is measured. When the measurement result meets the preset similarity condition, the newly arriving pulse clusters are determined to be delayed redundant pulses and interception is performed to generate a denoised discrete pulse data stream. Step 4: Convert the denoised discrete pulse data stream into a two-dimensional topological time feature map; Step 5: Input the two-dimensional topological time feature map into a deep cascaded spiking neural network for classification, and generate an early warning signal based on the classification result.

2. The method according to claim 1, characterized in that, In step two, the instantaneous pulse excitation density rate of the frequency channel address is calculated by setting a timing sliding monitoring window, and the length of the timing sliding monitoring window ranges from 1 microsecond to 5000 microseconds; the instantaneous pulse excitation density rate exceeding the preset density threshold means that the instantaneous pulse excitation density rate of one frequency channel address or multiple frequency-adjacent associated frequency channel addresses exceeds the preset density threshold within the time length of the timing sliding monitoring window.

3. The method according to claim 1, characterized in that, The channel-level dynamic refractory period mask window adopts an asymmetric mask adaptive shrinkage strategy based on the cross-band acoustic energy dissipation gradient. Specifically, based on the frequency band height of each frequency channel address contained in the original frequency envelope baseline vector of the direct wave, a set of frequency-related mask duration parameters are generated. According to the parameters, a channel-level dynamic refractory period mask window with an effective period shorter than the estimated reverberation time is set for frequency channel addresses with frequencies higher than a preset frequency threshold. At the same time, a channel-level dynamic refractory period mask window with an effective period equal to the estimated reverberation time is set for frequency channel addresses with frequencies lower than the preset frequency threshold.

4. The method according to claim 1, characterized in that, In step three, the similarity measurement is calculated using cosine similarity; the measurement result satisfies the preset similarity condition, specifically meaning that the result of the similarity measurement is greater than or equal to the preset similarity threshold. At this time, the newly arrived pulse cluster is determined to be a delayed redundant pulse and interception and discarding are performed.

5. The method according to claim 1, characterized in that, The method further includes a step of implementing adaptive decay suppression of network membrane potential between step four and step five. This step includes: calculating the number of pulses in the denoised discrete pulse data stream whose excitation intensity is lower than a set background noise energy threshold to assess the residual background noise level; and adaptively adjusting the membrane potential leakage integral constant of neurons in the deep cascaded spiking neural network according to the background noise level. When the background noise level is detected to be greater than a preset background noise threshold, the baseline leakage rate of neuron membrane potential is increased.

6. The method according to claim 1, characterized in that, In step four, the denoised discrete pulse data stream is converted into the two-dimensional topological time feature map by an exponential smooth decay algorithm based on the pulse arrival time.

7. The method according to claim 1, characterized in that, In step three, the estimated reverberation duration is set by calculating the estimated reverberation duration using the formula T_reverb = (2d / v) * k, where d is the distance between the sensor and the nearest acoustic reflector in the scene, v is the speed of sound propagation in the spatial medium, k is the environmental diffraction redundancy ratio coefficient, and T_reverb is the estimated reverberation duration.

8. The method according to claim 1, characterized in that, The deep cascaded spiking neural network is either a spiking convolutional neural architecture based on leakage integral and firing mechanism, or a spiking neural network architecture with a time-dependent synaptic feature self-attention mechanism.

9. The method according to claim 1, characterized in that, In step five, the early warning signal includes the classification result and the global reference timestamp of the first wave occurrence.

10. A reverberation-resistant processing system based on neuromorphic auditory events for abnormal voiceprint recognition in security scenarios, characterized in that, include: The purely asynchronous auditory acquisition module is used to acquire a sequence of purely asynchronous auditory discrete events and parse each discrete event in the sequence of purely asynchronous auditory discrete events into a digital tuple data packet containing a timestamp, polarity information and frequency channel address; The reference feature extraction module is used to calculate the instantaneous pulse excitation density rate of the frequency channel address. When the instantaneous pulse excitation density rate exceeds a preset density threshold, it is determined that a direct real original pulse cluster induced by the direct original sound wave has been captured. The arrival time of the energy spike of the direct-to-real original pulse cluster is set as the global reference timestamp of the first wave occurrence. Its spatial distribution in the frequency channel address dimension is extracted to determine the active channel set, and the original frequency envelope baseline vector of the direct wave is generated. The mask-side suppression filtering module is used to determine the estimated reverberation duration based on the physical characteristics of the deployment scenario; Starting from the global reference timestamp of the first wave occurrence, a channel-level dynamic refractory period mask window with an effective duration determined based on the estimated reverberation time is applied to the activated channel set. The similarity between the newly arriving pulse clusters within the effective period of the channel-level dynamic refractory period mask window and the original frequency envelope baseline vector of the direct wave is measured. When the measurement result meets the preset similarity condition, the newly arriving pulse clusters are determined to be delayed redundant pulses and interception is performed to generate a denoised discrete pulse data stream. An adaptive two-dimensional topology and noise reduction reconstruction module is used to convert the denoised discrete pulse data stream into a two-dimensional topological time feature map; The cascaded spiking neural network classification and early warning module is used to input the two-dimensional topological time feature map into a deep cascaded spiking neural network for classification, and generate an early warning signal based on the classification result.