Method and system for fast Fourier transform based on pulsed neuron network
Through analog-to-digital conversion, noise reduction and enhanced preprocessing, a pulsed neuron network is built, and sparseness analysis and alternative gradient optimization are used to solve the problems of wasted computing resources and low energy efficiency in sparse signal processing, realizing efficient and low-power Fourier transform.
Patent Information
- Application Number
- CN202510448842.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-04-10
AI Technical Summary
The traditional fast Fourier transform based on pulsed neuron networks has problems such as wasting computing resources and inefficient hardware when processing sparse signals. Especially in embedded devices, redundant calculations lead to high power consumption and the sparseness of signals cannot be automatically identified and utilized.
Through analog-to-digital conversion, noise reduction and enhanced preprocessing, a pulsed neuron network is constructed, sparse areas are identified using sparseness analysis, and the time-driven characteristics of the pulsed neuron network are expanded into static calculation diagrams. Custom operators are used to simulate the alternative gradient of pulse activation, and weight quantization and synaptic pruning optimization are performed.
Significantly reduces redundant computing load, improves computing efficiency and energy efficiency, is suitable for low-power edge computing scenarios, reduces multiplication operations by about 90%, and improves throughput.
Smart Images

Figure CN119939234B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital signal processing and artificial intelligence chips, and particularly relates to a method and system for fast Fourier transform based on a spiking neuron network. Background Art
[0002] Signal processing is a general term for the processing process of various types of electrical signals according to various expected purposes and requirements. The processing of analog signals is called analog signal processing, and the processing of digital signals is called digital signal processing. In actual signal processing, the frequency-domain energy distribution of most signals such as speech, bioelectric signals, and mechanical vibrations shows high sparsity, that is, the effective information of the signal is concentrated in a few frequency bands, and the energy of the remaining frequency bands can be ignored. Taking the speech signal as an example, its energy is mainly distributed in the narrow band corresponding to the fundamental frequency and harmonic components, while the energy of high-frequency noise and low-frequency environmental interference accounts for less than 10%. Similarly, in electroencephalogram signals, the characteristic high-frequency oscillation during epileptic seizures only occupies the local frequency band of 30Hz - 100Hz. For signal processing, the commonly used technology is fast Fourier transform (FFT) based on a spiking neuron network. The traditional algorithm for fast Fourier transform based on a spiking neuron network requires equal-weight calculation for all frequency points of the signal regardless of its actual energy contribution. This mode of non-discriminatory calculation in the entire frequency band leads to waste of computing resources and low hardware energy efficiency; the waste of computing resources is mainly reflected in: for a signal with a length of N, the traditional fast Fourier transform based on a spiking neuron network needs to perform complex multiplication and addition operations. Even if the signal has only k significant frequency bands, the computational complexity cannot be reduced; for example, when processing a 4096-point speech signal, the traditional fast Fourier transform based on a spiking neuron network needs to complete 24,576 complex multiplications, while the actual effective frequency band only accounts for about 20%, meaning that more than 19,000 multiplication operations are consumed on irrelevant frequency bands; the low hardware energy efficiency is mainly reflected in: in embedded devices such as hearing aids and wearable sensors, redundant calculations directly increase power consumption; actual measurements show that the energy consumption of the traditional fast Fourier transform based on a spiking neuron network for processing 1 second of speech signal on an ARM Cortex-M4 processor is 85mJ, and more than 60% of the power consumption is used for calculating low-energy frequency bands.
[0003] Although the traditional fast Fourier transform based on spiking neural networks widely satisfies frequency domain analysis, its processing method is not efficient for some application scenarios. Especially when the key frequency points cannot be accurately captured, it is easy to cause unnecessary power consumption waste. Even if the useful information of the signal only exists in a small part of the data, in order to obtain the complete spectrum, it is still necessary to calculate all N data points, which increases the computational burden. Especially many signals in nature (speech, images) are sparse in the frequency domain, that is, only a few frequency components have significant energy. The fast Fourier transform based on spiking neural networks cannot automatically identify and utilize this sparsity, but treats all frequency bins equally, resulting in problems such as low computational efficiency, high latency, and large resource occupancy. Summary of the Invention
[0004] To achieve the above object, the present invention adopts the following technical solutions:
[0005] On the one hand, the present invention provides a method for fast Fourier transform based on spiking neural networks, comprising the following steps:
[0006] Collect an analog signal through an analog-to-digital converter at a configured sampling rate parameter, convert the analog signal into a digital signal, and the converted digital signal enters a field-programmable gate array after passing through an anti-aliasing filter; the digital signal entering the field-programmable gate array is preprocessed for noise reduction and enhancement to obtain a preprocessed digital signal;
[0007] Construct a spiking neural network according to the number of points of the fast Fourier transform based on spiking neural networks; use sparsity analysis to identify the sparse regions of the preprocessed digital signal and neuron activities;
[0008] Deploy the spiking neural network on a neural network processor, expand the time-driven characteristic of the spiking neural network into a static computational graph, and simulate the surrogate gradient of pulse activation through a custom operator; use the neural network processor offline model conversion tool for weight quantization and synaptic pruning optimization.
[0009] In an optional implementation manner, the process of obtaining the preprocessed digital signal comprises the following steps:
[0010] The multi-channel interface is connected to the memory of the neural network processor, and the analog signal is collected in parallel through the multi-channel interface; the digital signal entering the field-programmable gate array is decomposed into a noise reduction sub-task and an enhancement sub-task;
[0011] The noise reduction sub-task decomposes the wavelet threshold denoising algorithm into multiple 3×3 convolution kernels and processes 8 channels of signals on the neural network processor carried by the wavelet threshold denoising algorithm; the enhancement sub-task uses a vector unit to perform normalization and dynamic range compression;
[0012] The digital signal after obtaining the noise reduction subtask and the enhancement subtask is input into the spiking neuron network.
[0013] In an optional implementation manner, the processing process of the noise reduction subtask and the enhancement subtask includes the following steps:
[0014] The input digital signal is segmented according to a time window, and each frame of digital signal is modulated by a complex adjustable basis function and decomposed into high-frequency and low-frequency components; according to the local transient energy gradient of the digital signal, the threshold intensity is dynamically adjusted to achieve discriminative suppression in an environment mixed with impulsive noise and steady-state noise;
[0015] The wavelet band-pass filtering is equivalent to a 3×3 reconfigurable convolution kernel array, and the parameters in the kernel are updated frame by frame according to the signal time-frequency ridge line characteristics. For continuous digital signals, a smoothing kernel is used; for transient pulse digital signals, it is switched to a differential kernel;
[0016] The vector processing unit carried by the neural network processor adopts a mixed-precision serial calculation stream. The input signal is block-divided for zero-phase shift filtering, the dynamic gain curve is calculated in the time domain, and the non-linear energy compression is realized by using the logarithmic Hilbert transform of the digital signal envelope.
[0017] In an optional implementation manner, the process of performing noise reduction and enhancement preprocessing further includes the following steps:
[0018] The input digital signal is regularly divided into 16×16 blocks to obtain data blocks; the data blocks are alternately stored in 4 independent memory banks by using memory bank interleaving; the SIMD parallel capability of the NEON instruction set is used, and a single VMLA instruction simultaneously completes 4 groups of complex multiplication operations, and data duplication loading is avoided through register reuse technology;
[0019] The scalar unit of the neural network processor adopts sliding window statistics to calculate the mean / standard deviation of the signal amplitude in the window in real time, accelerates the standard deviation calculation by using the hardware square root unit, generates a dynamic threshold and synchronizes it through the broadcast bus; the tensor unit parallelly loads 16×16 data blocks, uses vector comparison instructions to perform threshold judgment, generates a sparse mask, executes BITPACK compression, and stores non-zero data and coordinates in CSC format;
[0020] A scalar-tensor double-buffer pipeline is constructed. The threshold calculation of the current frame is parallel to the sparse processing of the previous frame, and data is automatically transferred through DMA. The threshold bus ensures calculation synchronization; the dynamic range of the signal is detected in real time, the FP16 calculation mode is switched for a small dynamic range, the all-zero mask triggers the calculation to skip, and the voltage and frequency are dynamically adjusted according to the load.
[0021] In an optional implementation manner, the process of constructing the spiking neuron network includes the following steps:
[0022] Construct a spiking neural network according to the number of points of the fast Fourier transform based on the spiking neuron network. Each layer contains a pair of leaky integrate-and-fire neurons, and each pair of leaky integrate-and-fire neurons processes the real and imaginary parts of a complex number for each group.
[0023] Decompose the rotation factor into real and imaginary parts, and define the weight matrix for each pair of neurons. After the input spike signal is integrated with weights, calculate the input current.
[0024] Butterfly span calculation, calculate the butterfly distance and the output of butterfly nodes in a certain layer. The last layer is spike frequency encoding, that is, count the average number of spikes in the output layer, and after normalization, obtain the frequency domain amplitude, thus completing the fast Fourier transform based on the spiking neuron network.
[0025] In an optional implementation, the process of each pair of leaky integrate-and-fire neurons processing the real and imaginary parts of a complex number for each group includes the following steps:
[0026] The spike output signal of the leaky integrate-and-fire neuron flows through the decoding module. Through the differential receiver at the dendrite input end, the spike cluster is split into signals and fed into the dendrite circuit array. The enable signal triggers the intelligent gate driver of the drive module to activate the time gate synchronization sequence.
[0027] The real part data stream is injected into the capacitor network of the dendrite circuit, and the weight is adaptively adjusted through the charge domain effect. The imaginary part path completes amplitude-phase correction inside the drive module.
[0028] The soma circuit executes, and the Schmitt trigger compares the RC integration values of the two paths in real time. When the vector amplitude exceeds the threshold of the programmable comparator, it triggers the transition latch to update the rotation phase parameter, and sends the encoded pulse to the output bus through the final stage driver.
[0029] In an optional implementation, the process of using sparsity analysis to identify the sparse regions of the preprocessed digital signal and neuron activity includes the following steps:
[0030] In the spiking neuron network, decompose the digital signal into sine waves of different frequencies through the fast Fourier transform based on the spiking neuron network to obtain the frequency domain representation of the signal. In the spiking neuron network, calculate the L1 norm of the spike signals of the input layer neurons to obtain the sparse metric value of the signal.
[0031] Through the sparsity analysis of the spike signals of the input layer neurons, extract the important features of the digital signal as the sparse regions.
[0032] Through threshold processing of the spike signals of the output layer neurons, find the neurons with sparsity higher than the threshold, and the corresponding spike signals of the neurons are the sparse regions of the digital signal.
[0033] In an alternative embodiment, a process of simulating the surrogate gradient of pulse activation through a custom operator includes the following steps:
[0034] When the membrane potential of a neuron exceeds the threshold, the dynamic scaling property of the hyperbolic tangent function is used as the surrogate gradient kernel. The hyperbolic tangent function forms a continuously differentiable saturation region near the pulse triggering threshold, and the radius of curvature thereof is negatively correlated with the refractory period characteristic of the neuron, such that the pseudo-derivative at the threshold crossing point can faithfully reflect the temporal firing pattern of biological neurons.
[0035] Introduce a time expansion factor to map the time-domain discrete events of the pulse sequence into virtual connection paths in a multi-dimensional tensor space; each pulse event generates a virtual wire with an attenuation factor along the time axis, and the virtual wires form dynamically conductive differential channels in the surrogate gradient field.
[0036] When a pulsed neuron is triggered, discrete coordinates in COO format are generated during forward propagation, and during the backpropagation stage, a surrogate gradient field under four-dimensional spatio-temporal constraints is reconstructed through the discrete coordinates. The surrogate gradient field forms a conjugate match with the partial differential of the pulse density function; threshold modulation similar to the photoelectric effect is used to dynamically adjust the amplitude of the surrogate gradient with the pulse interphase, matching the asynchronous computing characteristics of the Cube Unit in the neural processor.
[0037] On the other hand, the present invention provides a system for fast Fourier transform based on a pulsed neuron network for implementing the method for fast Fourier transform based on a pulsed neuron network. The system for fast Fourier transform based on a pulsed neuron network includes:
[0038] A signal processing module for collecting an analog signal through an analog-to-digital converter at a configured sampling rate parameter, converting the analog signal into a digital signal, and after the converted digital signal passes through an anti-aliasing filter, entering a field programmable gate array; the digital signal entering the field programmable gate array is subjected to noise reduction and enhancement preprocessing to obtain a preprocessed digital signal;
[0039] A network construction module for constructing a pulsed neuron network according to the number of points of the fast Fourier transform based on a pulsed neuron network. Each layer of neurons in the pulsed neuron network receives circular pulse inputs, updates the membrane potential and triggers a pulse through weighted sum calculation, and outputs the pulse density of the layer, which is converted into a frequency domain assignment through integration; sparsity analysis is used to identify the sparse regions of the preprocessed digital signal and neuron activities; the pulsed neuron network includes a network structure design unit, a neuron model selection unit, and a connection weight setting unit;
[0040] The model conversion module is used to deploy the spiking neuron network on the neural network processor, expand the time-driven characteristics of the spiking neuron network into a static computational graph, and simulate the surrogate gradient of spiking activation through custom operators; use the neural network processor offline model conversion tool for weight quantization and synaptic pruning optimization; when deploying, adopt the dynamic sparse acceleration ability of the neural network processor, encode the spiking events into a COO format sparse tensor, and directly trigger the sparse matrix multiplication of the Cube Unit.
[0041] On the other hand, the present invention provides an application of the method for fast Fourier transform based on the spiking neuron network in a neural network processor, including: the neural network processor is an AI chip.
[0042] The present invention collects continuous time-domain signals through an analog-to-digital converter (ADC) according to set sampling rate parameters; the output signal passes through an anti-aliasing filter to suppress high-frequency noise, ensuring that the subsequent processing frequency band is limited within an effective range; the digital signal is denoised and enhanced in a field-programmable gate array (FPGA) to optimize the input quality of subsequent frequency-domain analysis. An SNN is constructed according to the number of FFT points, and its time encoding characteristic is used to map the time-domain signal into a spike density, and the frequency-domain amplitude is output through integration; combined with the spatio-temporal sparsity of the input signal and neuron activation, the redundant calculation load is dynamically reduced; the SNN layer (such as the input layer mapping time window, hidden layer non-linear transformation) is determined, and the (LIF) model is used to realize the dynamic update of the membrane potential and optimize the spike transmission efficiency. The time-driven logic of the SNN is expanded into a static data flow graph, the non-differentiable problem of spikes is solved through the surrogate gradient, and the ONNX inference framework is compatible; the spike events are encoded into a COO sparse tensor, triggering the sparse matrix multiplication of the Cube Unit of the NPU (skipping zero-value calculations), significantly improving the throughput. Description of the Drawings
[0043] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. They are used to explain the present invention together with the embodiments of the present invention, and do not constitute a limitation to the present invention. In the drawings:
[0044] Figure 1 It is a flowchart of the method for fast Fourier transform based on the spiking neuron network provided in Embodiment 1 of the present invention;
[0045] Figure 2 It is a process diagram of obtaining the preprocessed digital signal provided in Embodiment 2 of the present invention;
[0046] Figure 3 It is a process diagram of constructing the spiking neuron network provided in Embodiment 3 of the present invention;
[0047] Figure 4It is a process diagram of simulating the surrogate gradient of pulse activation through a custom operator provided in Embodiment 4 of the present invention;
[0048] Figure 5 It is a system block diagram of fast Fourier transform based on a spiking neuron network provided in Embodiment 9 of the present invention;
[0049] Figure 6 It is a schematic application diagram of the method of fast Fourier transform based on a spiking neuron network in a neural network processor in Embodiment 10 of the present invention;
[0050] Figure 7 It is a schematic diagram of the LIF neuron circuit model in Embodiment 10 of the present invention;
[0051] Figure 8 It is a block diagram of the electronic device provided by the present invention;
[0052] Figure 9 It is a block diagram of the computer-readable storage medium provided by the present invention. Detailed implementation manners
[0053] Next, the technical solutions in the embodiments of the present invention will be described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0054] Hereinafter, terms such as "first" and "second" are only for convenience of description and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present invention, unless otherwise stated, the meaning of "a plurality" is two or more.
[0055] In the present invention, unless otherwise clearly specified and defined, the term "connection" should be understood in a broad sense. For example, "connection" can be a fixed mechanical connection, a detachable mechanical connection, or integrated; or, "connection" can be a direct connection, or an indirect connection through an intermediate medium. In addition, unless otherwise clearly specified and defined, the term "coupling" should be understood in a broad sense. For example, "coupling" can be a direct electrical connection. For example, physical contact and electrical conduction occur between two components, and it can also be understood that different components in a circuit structure are electrically connected through an entity line such as a copper foil or a wire of a printed circuit board (PCB) that can transmit electrical signals for the transmission of electrical signals; or, "coupling" can be an indirect electrical connection between two components through an intermediate medium; or, "coupling" can be an electrical connection between two components in a non-contact manner, such as an electrical connection between two components in a capacitive coupling manner for the transmission of electrical signals.
[0056] In the embodiments of the present invention, orientation terms such as "upper", "lower", "left", "right", etc. can include but are not limited to being defined relative to the schematic placement of components in the drawings. It should be understood that these directional terms can be relative concepts, and they are used for relative description and clarification, and they can change accordingly with the change of the orientation of the components in the drawings.
[0057] Embodiment 1:
[0058] As Figure 1 shown, the embodiments of the present invention provide a method for fast Fourier transform based on a pulsed neuron network, including the following steps:
[0059] Step S100: Collect an analog signal through an analog-to-digital converter at a configured sampling rate parameter, convert the analog signal into a digital signal, and the converted digital signal enters a field-programmable gate array after passing through an anti-aliasing filter; the digital signal entering the field-programmable gate array is preprocessed for noise reduction and enhancement to obtain a preprocessed digital signal;
[0060] Step S200: Construct a pulsed neuron network according to the number of points of the fast Fourier transform based on the pulsed neuron network. Each layer of neurons in the pulsed neuron network receives a spherical pulse input, updates the membrane potential and triggers a pulse through weighted sum calculation, outputs the pulse density of the layer, and is converted into a frequency domain assignment through integration; use sparsity analysis to identify the sparse regions of the preprocessed digital signal and neuron activity; the pulsed neuron network includes a network structure design unit, a neuron model selection unit, and a connection weight setting unit;
[0061] Step S300: Deploy the spiking neuron network on the neural network processor, expand the time-driven characteristics of the spiking neuron network into a static computational graph, and simulate the surrogate gradient of spiking activation through custom operators; use the neural network processor offline model conversion tool for weight quantization and synaptic pruning optimization; when deploying, adopt the dynamic sparse acceleration ability of the neural network processor, encode the spiking events into a COO format sparse tensor, and directly trigger the sparse matrix multiplication of the Cube Unit.
[0062] In the above embodiments, in step S100, the analog-to-digital converter (ADC) is used to collect continuous time-domain signals according to the set sampling rate parameters; the output signal passes through an anti-aliasing filter to suppress high-frequency noise, ensuring that the subsequent processing frequency band is limited within the effective range; the digital signal is denoised and enhanced in the field-programmable gate array (FPGA) to optimize the input quality of subsequent frequency-domain analysis. In step S200, an SNN is constructed according to the number of FFT points, and its time encoding characteristic is used to map the time-domain signal into a spike density, and the frequency-domain amplitude is output through integration; combined with the spatio-temporal sparsity of the input signal and neuron activation, the redundant calculation load is dynamically reduced; the SNN layer (such as the input layer mapping time window, hidden layer non-linear transformation) is determined, and the (LIF) model is used to realize the dynamic update of the membrane potential and optimize the spike transmission efficiency. In step S300, the time-driven logic of the SNN is expanded into a static data flow graph, the non-differentiable problem of spikes is solved through the surrogate gradient, and the ONNX inference framework is compatible; the spiking events are encoded into a COO sparse tensor, triggering the sparse matrix multiplication of the Cube Unit of the NPU (skipping zero-value calculations), significantly improving the throughput.
[0063] Embodiment 2:
[0064] As Figure 2 shown, on the basis of Embodiment 1, the process of obtaining the preprocessed digital signal in step S100 provided by the embodiment of the present invention includes the following steps:
[0065] Step S101: The multi-channel interface is connected to the memory of the neural network processor, and the analog signal is collected in parallel through the multi-channel interface; the digital signal entering the field-programmable gate array is decomposed into a denoising subtask and an enhancement subtask;
[0066] Step S102: The denoising subtask decomposes the wavelet threshold denoising algorithm into multiple 3×3 convolution kernels and processes 8 channels of signals on the neural network processor carried by the wavelet threshold denoising algorithm; the enhancement subtask uses the vector unit to perform normalization and dynamic range compression;
[0067] Step S103: Obtain the digital signal after the denoising subtask and the enhancement subtask, and input it into the spiking neuron network.
[0068] In the above embodiments, for step S101 of signal acquisition and subtask decomposition, a high-speed analog-to-digital converter and a time-division multiplexing mechanism are adopted to ensure synchronous sampling of multi-channel analog signals and reduce crosstalk and jitter; for step S102 of subtask parallel processing, each core processes 2 signals, with a total of 8 signals in parallel, reducing memory access latency; enhancing the vectorized execution of subtasks, and normalizing single-instruction to complete multi-data calculations; dynamic range compression reduces the non-linear distortion of high-dynamic range data.
[0069] In this embodiment, ±10V analog signals are collected through a 16-channel differential analog-to-digital converter ADC, the sampling rate is configured as 1MSPS, and after passing through an anti-aliasing filter, it is input into the FPGA. At the same time, it supports the parallel data stream of the LVDS interface and is directly connected to the Ascend 310B development board through an FMC connector. For high-speed pulse sequences (such as the output of a neuromorphic camera), they are input through a dedicated AXI-Stream interface, and timestamp compression coding is used to reduce bandwidth occupancy. For data transmission, a PCIe 3.0 x4 interface or the dedicated AXI bus of the Ascend chip is used to directly write the original data into the LPDDR4X memory pool of the Ascend chip, bypassing the CPU to reduce latency.
[0070] Embodiment 3:
[0071] Based on Embodiment 2, the processing procedures of the noise reduction subtask and the enhancement subtask in step S102 provided by the embodiment of the present invention include the following steps:
[0072] Step S1021: The input digital signal is segmented by time window, and each frame of digital signal is modulated by a complex adjustable basis function and decomposed into high-frequency and low-frequency components; according to the local transient energy gradient of the digital signal, the threshold intensity is dynamically adjusted to achieve discriminative suppression in an environment of mixed impulsive noise and steady-state noise;
[0073] Step S1022: Equivalent the wavelet band-pass filter to a 3×3 reconfigurable convolution kernel array, and the parameters inside the kernel are updated frame by frame according to the time-frequency ridge line characteristics of the signal. For continuous digital signals, a smoothing kernel is used; for transient pulse digital signals, it is switched to a differential kernel;
[0074] Step S1023: The vector processing unit carried by the neural network processor adopts a mixed-precision serial calculation stream. The input signal is block-processed for zero-phase shift filtering, the dynamic gain curve is calculated in the time domain, and non-linear energy compression is achieved by using the logarithmic Hilbert transform of the digital signal envelope.
[0075] In the above embodiments, the complex adjustable basis function modulation in step S1021 breaks through the limitation of the fixed wavelet basis. By dynamically adjusting the scale factor and phase shift of the basis function, the optimal sparse representation of the digital signal in the time-frequency two-dimensional plane is realized, improving the frequency-domain resolution and the ability to capture transient components. The transient energy gradient threshold dynamically generates a threshold curve based on the local microstructure characteristics of the digital signal, suppressing the steady-state background noise while retaining the pulse signal. In step S1022, the reconfigurable convolution kernel array realizes the real-time loading of the filter kernel parameters through hardware programmable logic, enabling a single set of convolution kernels to have both smoothing and differential functions, and avoiding the phase distortion problem of cascading multiple-stage filters in traditional methods. The parameter update guided by the time-frequency ridge line feature uses the instantaneous frequency derivative of the signal as the basis for adjusting the kernel weights, enabling the spatial domain convolution operation to have the ability of time-frequency joint analysis. In step S1023, zero-phase shift filtering eliminates the group delay of traditional causal filtering through bidirectional recursive filtering, ensuring the time-domain alignment of the signal, and is applicable to the multi-sensor data fusion scenario. The logarithmic Hilbert transform compression maps the envelope energy of the digital signal to the logarithmic domain and performs analytic signal reconstruction, realizing equivalent dynamic range compression on the premise of avoiding frequency-domain transformation.
[0076] Embodiment 4:
[0077] Based on Embodiment 1, the process of performing noise reduction and enhancement preprocessing in step S100 provided by the embodiment of the present invention further includes the following steps:
[0078] Step S104: Regularly divide the input digital signal into 16×16 blocks to obtain data blocks; alternately store the data blocks into 4 independent memory banks using memory bank interleaving; use the SIMD parallel ability of the NEON instruction set, and a single VMLA instruction simultaneously completes 4 groups of complex multiplication operations. By using the register reuse technology to avoid repeated data loading, the calculation speed of the rotation factor is improved;
[0079] Step S105: The scalar unit of the neural network processor uses a sliding window statistic to calculate the mean / standard deviation of the signal amplitude within the window in real time, accelerates the standard deviation calculation using the hardware square root unit, generates a dynamic threshold and synchronizes it through the broadcast bus; the tensor unit parallelly loads 16×16 data blocks, uses vector comparison instructions to perform threshold judgment, generates a sparse mask, performs BITPACK compression, and stores the non-zero data and coordinates in the CSC format;
[0080] Step S106: Construct a scalar-tensor double-buffer pipeline, parallelize the threshold calculation of the current frame and the sparse processing of the previous frame, realize automatic data transfer through DMA, and ensure calculation synchronization through the threshold bus; detect the dynamic range of the signal in real time, switch to the FP16 calculation mode for small dynamic ranges, trigger calculation skipping for all-zero masks, and dynamically adjust the voltage and frequency according to the load.
[0081] In the above embodiments, step S104 performs parallel FFT preprocessing based on NEON to solve the memory access bottleneck of traditional FFT, and realizes full pipelining of data supply and computing units; it provides high-quality spectral data with time domain alignment for sparsification processing. Step S105 performs dynamic sparsification encoding to adaptively eliminate noise / redundant spectral components, compressing the data volume to 9.3% - 22.1% of the original size and saving 62% of the memory bandwidth. Step S106 performs system-level dynamic optimization to achieve a balance in the utilization rate of scalar-tensor units, reducing the power consumption performance to 37% - 52% of the fixed mode and ensuring real-time requirements.
[0082] Embodiment 5:
[0083] As Figure 3 shown, on the basis of Embodiment 1, the process of constructing a spiking neuron network in step S200 provided by the embodiment of the present invention includes the following steps:
[0084] Step S201: Construct a spiking neural network according to the number of points of the fast Fourier transform based on the spiking neuron network. Each layer contains a pair of leaky integrate-and-fire neurons, and the pair of leaky integrate-and-fire neurons processes the real and imaginary parts of each complex number;
[0085] Step S202: Decompose the rotation factor into a real part and an imaginary part, and define the weight matrix of each neuron pair; after the input pulse signal is integrated by the weight, calculate the input current;
[0086] Step S203: Calculate the butterfly span, calculate the butterfly distance and the output of the butterfly node of a certain layer; the last layer is pulse frequency encoding, that is, count the average number of pulses in the output layer, and obtain the frequency domain amplitude after normalization, completing the fast Fourier transform based on the spiking neuron network.
[0087] Among them, construct layer SNN according to the number of points N of the fast Fourier transform based on the spiking neuron network. The number of neurons in each layer is N / 2, corresponding to the number of butterfly operation nodes; for example, an 8-point FFT requires a 3-layer structure, and each layer contains 4 pairs of LIF neurons (processing the real and imaginary parts of complex operations). LIF parameter settings: Configure the membrane time constant = 10ms, the membrane potential threshold V th = 1.0mV, and the membrane potential update formula is:
[0088]
[0089] Among them represents the membrane potential at the current moment , in millivolts (mV); represents the membrane potential at the previous moment ; Indicates the current time The input current, with the unit of ampere (A); Indicates the membrane time constant, representing the response speed of the membrane potential to the change of the input current, which is 10 ms here;
[0090] According to the FFT rotation factor Set the synaptic weights; the weight matrix of each pair of neurons (real part, imaginary part) is split into:
[0091] ,
[0092]
[0093] where Represents the complex exponential form, used for frequency-domain analysis, k represents the frequency index, and N represents the total number of samples; Represents the connection weight from the real part neuron to the real part neuron; Represents the connection weight from the real part neuron to the imaginary part neuron; Represents the connection weight from the imaginary part neuron to the real part neuron; Represents the connection weight from the imaginary part neuron to the imaginary part neuron; Represents the rotation angle in the complex plane (unit: radian).
[0094] Each layer of neurons receives the pulse input from the previous layer and calculates through the weighted sum , updates the membrane potential and triggers a pulse. For example, the output of the th butterfly node in the mth layer is: . Where Δ represents the butterfly span of the current layer; m represents the layer index of the current neural network; Represents the th butterfly operation node in the current layer; Represents the synaptic weight matrix; LIF(·) represents the leaky integrate-and-fire model; Represents the pulse stream received by the current neuron; Represents the butterfly connection input across Δ neurons; Represents the output of the th butterfly node in the mth layer; the pulse density of the final output layer (number of pulses per unit time) is converted to the frequency-domain amplitude through integration to complete the FFT calculation.
[0095] In the above embodiments, step S201 constructs a spiking neural network (SNN) based on the number of points of the fast Fourier transform (FFT) of the spiking neuron network, and realizes parallel processing of complex signals by hierarchically deploying leaky integrate-and-fire neuron pairs. Among them, each group of neuron pairs processes the real part and the imaginary part of the complex number respectively, ensuring the completeness of complex number operations in the spiking domain; it provides a hardware-implementable spiking computing infrastructure for the FFT of the spiking neuron network. Step S202 decomposes the rotation factor of the FFT into the real part of the corresponding even-symmetric component and the imaginary part of the corresponding odd-symmetric component , and maps them to the weight matrix of the neuron pair; the input spiking signal generates an input current as a spiking sequence after weighted integration, driving the membrane potential dynamics of the LIF neuron; it converts the frequency-domain operation into a synaptic plasticity problem of spiking spatio-temporal coding. Step S203 calculates the span of the butterfly unit layer by layer, realizes in-situ calculation by adjusting the spiking routing connection between neurons, and reduces the memory access overhead; at the output layer, the average firing rate of neurons is statistically calculated, and the frequency-domain amplitude is obtained after normalization; it completes the non-linear mapping from the time-domain spiking sequence to the frequency-domain amplitude, satisfying the energy conservation characteristic of the FFT.
[0096] In summary, this embodiment converts the complex multiplication and accumulation operation of the FFT into a spatio-temporal event-driven calculation of the spiking neural network, which is applicable to low-power edge computing scenarios; it realizes integration in the analog domain through the leakage-threshold characteristic of the LIF neuron, reducing about 90% of the multiplication operations compared with the digital FFT.
[0097] Embodiment 6:
[0098] Based on Embodiment 5, the process of each group of leaky integrate-and-fire neuron pairs in step S201 provided by the embodiment of the present invention for processing the real part and the imaginary part of the complex number includes the following steps:
[0099] Step S2011: The spiking output signal of the leaky integrate-and-fire neuron flows through the decoding module. Through the differential receiver at the dendrite input end, the spiking cluster is split into signals and fed into the dendrite circuit array; the enable signal triggers the intelligent gate driver of the driving module to activate the time gate synchronization sequence;
[0100] Step S2012: The real part data stream is injected into the capacitive network of the dendrite circuit, and the weight is adaptively adjusted through the charge domain effect; the imaginary part path completes amplitude-phase correction inside the driving module;
[0101] Step S2013: The soma circuit executes, and the Schmidt trigger compares the RC integration values of the two paths in real time. When the vector amplitude exceeds the threshold of the programmable comparator, it triggers the transition latch to update the rotation phase parameter, and sends an encoded pulse to the output bus through the final driver.
[0102] In the above embodiments, in step S2011 of pulse decoding and synchronous activation, the pulse output signal of the leakage integration-emitting neuron first undergoes processing by the decoding module; through the differential amplifier at the dendritic input end, the pulse signal is split into multiple sub-signals, which are respectively fed into the dendritic circuit array, achieving the spatial distribution of the pulse signal; meanwhile, the enable signal triggers the intelligent gate driver in the driving module to activate the time grid synchronization sequence, ensuring the timing control of signal processing, enabling each sub-signal to be processed synchronously, and guaranteeing the overall performance of the system. In step S2012 of real-part data injection and imaginary-part amplitude-phase correction, during the real-part data stream injection phase, the data signal is injected into the capacitive network of the dendritic circuit; the weights are adaptively adjusted through the charge-domain effect, dynamically adjusting the weights of the signals to adapt to different input signal characteristics; the imaginary-part path completes the amplitude-phase correction inside the driving module. Amplitude-phase correction is to correct the amplitude and phase of the imaginary-part signal to eliminate distortion and delay during signal transmission and improve the accuracy of signal processing. In step S2013 of the soma circuit execution and output pulse coding, the soma circuit is responsible for performing the final processing of the signal. The Schmitt trigger continuously compares the RC integration values of the real-part and imaginary-part paths. When the vector amplitude exceeds the threshold of the programmable comparator, it triggers the transition latch to update the rotation phase parameter; the update of the rotation phase parameter realizes the dynamic adjustment of the signal vector, enabling the signal vector to always maintain the optimal state; through the final-stage driver, the encoded pulse signal is sent to the output bus to complete the processing and output of the signal.
[0103] Embodiment 7:
[0104] Based on Embodiment 1, the process of using sparsity analysis to identify the sparse regions of the preprocessed digital signal and neuron activity in step S200 provided by the embodiment of the present invention includes the following steps:
[0105] Step S204: In the pulsed neuron network, the digital signal is decomposed into sine waves of different frequencies through fast Fourier transform based on the pulsed neuron network to obtain the frequency-domain representation of the signal; in the pulsed neuron network, by calculating the L1 norm of the pulse signals of the input-layer neurons, the sparse metric value of the signal is obtained;
[0106] Step S205: Through sparsity analysis of the pulse signals of the input-layer neurons, the important features of the digital signal are extracted as the sparse regions;
[0107] Step S206: By performing threshold processing on the pulse signals of the output-layer neurons, the neurons with a sparsity higher than the threshold are found, and the pulse signals corresponding to the neurons are the sparse regions of the digital signal.
[0108] In the above embodiments, in step S204, signal decomposition and sparsity measurement, in the spiking neuron network, the digital signal in the time domain is decomposed into sine waves of different frequencies through the fast Fourier transform based on the spiking neuron network to obtain the frequency-domain representation of the signal, and the digital signal is converted from the time domain to the frequency domain, which is convenient for analyzing the frequency characteristics and sparsity of the digital signal; by calculating the L1 norm of the spike signals of the input-layer neurons, the sparsity measurement value of the signal is obtained; the purpose of calculating the L1 norm is to measure the sparsity degree of the digital signal, and the sparse signal will have a smaller value under the L1 norm; by calculating the L1 norm, the sparsity of the digital signal can be quantified. In step S205, sparsity analysis and feature extraction, by performing sparsity analysis on the spike signals of the input-layer neurons, the important features of the digital signal are extracted, the key information and non-redundant parts in the digital signal are identified, providing a basis for sparse region localization; sparsity analysis usually involves analyzing the energy distribution of the digital signal to find out which frequency components the energy is mainly concentrated on, and the frequency components often correspond to the important features of the digital signal. By extracting these features, effective sparse coding of the digital signal is performed. In step S206, threshold processing and sparse region localization, by performing threshold processing on the spike signals of the output-layer neurons, the neurons with sparsity higher than the threshold are found, and the most significant part of the digital signal, that is, the sparse region, is identified; the spike signals corresponding to the neurons with sparsity higher than the threshold are the sparse regions of the digital signal. These sparse regions contain the core information of the signal and are of great significance for signal representation and processing; threshold processing is an effective sparsity quantization method. By setting a reasonable threshold, the sparse part and non-sparse part in the digital signal are distinguished, facilitating digital signal processing and analysis.
[0109] In summary, this embodiment completes the whole process from signal decomposition, sparsity measurement, feature extraction to sparse region localization, providing technical support for identifying and processing the sparse regions of digital signals; effectively extracting the key information of the signal, reducing redundancy, and improving the efficiency and accuracy of signal processing.
[0110] Embodiment 8:
[0111] As Figure 4 shown, on the basis of Embodiment 1, in step S300 provided by the embodiment of the present invention, the process of simulating the surrogate gradient of spike activation through a custom operator includes the following steps:
[0112] Step S301: When the membrane potential of the neuron exceeds the threshold, the dynamic scaling characteristic of the hyperbolic tangent function is used as the surrogate gradient kernel. The hyperbolic tangent function forms a continuously differentiable saturation region near the spike trigger threshold, and its radius of curvature is negatively correlated with the refractory period characteristic of the neuron, so that the pseudo-derivative at the threshold crossing point can faithfully reflect the temporal firing pattern of biological neurons;
[0113] Step S302: Introduce a time expansion factor to map the time-domain discrete events of the pulse sequence into virtual connection paths in a multi-dimensional tensor space; each pulse event generates an imaginary wire with an attenuation factor along the time axis, and the imaginary wire forms a dynamically conducting differential channel in the alternative gradient field;
[0114] Step S303: When the pulse neuron is triggered, discrete coordinates in COO format are generated during forward propagation, while during the backpropagation stage, an alternative gradient field under four-dimensional spatio-temporal constraints is reconstructed through the discrete coordinates. The conjugate matching is formed by the partial differential of the alternative gradient field and the pulse density function; threshold modulation similar to the photoelectric effect is adopted to make the amplitude of the alternative gradient dynamically adjusted with the pulse interphase, matching the asynchronous computing characteristics of the Cube Unit in the neural processor.
[0115] In the above embodiments, in step S301, the dynamic scaling characteristic of the hyperbolic tangent function converts the non-differentiable problem of discrete pulses into a continuously differentiable pseudo-gradient problem, ensuring the feasibility of backpropagation; through the negative correlation between the radius of curvature and the refractory period, the inherent time characteristics of real neurons (such as repolarization delay) are simulated, making the gradient calculation conform to the firing dynamics law; a saturation region is formed near the triggering threshold, which not only maintains the discreteness of pulse events but also retains the continuity of the gradient flow, solving the gradient disappearance problem of traditional step functions during training. In step S302, through the time expansion factor, discrete time-domain pulse events are projected into a multi-dimensional tensor space to form a computable virtual topological structure; the imaginary wire with an attenuation factor simulates the time-dependence of biological synapses (such as short-term plasticity), ensuring that the gradient is dynamically adjusted with the pulse interval and enhancing the ability to extract temporal features; the conduction path of the alternative gradient field is updated instantaneously according to pulse events, avoiding redundant calculations of traditional static computational graphs and adapting to asynchronous sparse pulse streams. In step S303, sparse coding of discrete coordinates in COO format is used to significantly reduce the storage and computational overhead of backpropagation, directly compatible with the sparse matrix multiplication unit of the brain-inspired chip; through the inverse mapping of discrete coordinates to spatio-temporal constraints, it is ensured that the gradient propagation path is strictly consistent with the forward propagation, avoiding error diffusion; the dynamic adjustment of the gradient amplitude with the pulse interphase simulates the adaptive change of synaptic weights with the firing frequency, realizing the collaborative optimization of hardware computing and biological models.
[0116] Embodiment 9:
[0117] As Figure 5 shown, based on Embodiments 1 - 8, the system for fast Fourier transform based on a pulse neuron network provided by the embodiment of the present invention includes:
[0118] The signal processing module 1 is used to collect analog signals through an analog-to-digital converter at a configured sampling rate parameter, convert the analog signals into digital signals, and the converted digital signals enter the field-programmable gate array after passing through an anti-aliasing filter; the digital signals entering the field-programmable gate array are preprocessed for noise reduction and enhancement to obtain preprocessed digital signals;
[0119] The network construction module 2 is used to construct a spiking neuron network according to the number of points of the fast Fourier transform based on the spiking neuron network. Each layer of neurons in the spiking neuron network receives a circle of spike inputs, updates the membrane potential and triggers spikes through weighted sum calculation, outputs the spike density of the layer, and converts it into a frequency domain assignment through integration; uses sparsity analysis to identify the sparse regions of the preprocessed digital signals and neuron activities; the spiking neuron network includes a network structure design unit, a neuron model selection unit, and a connection weight setting unit;
[0120] The model conversion module 3 is used to deploy the spiking neuron network on a neural network processor, expand the time-driven characteristics of the spiking neuron network into a static computational graph, and simulate the surrogate gradient of spike activation through a custom operator; use the offline model conversion tool of the neural network processor for weight quantization and synaptic pruning optimization; when deploying, adopt the dynamic sparse acceleration ability of the neural network processor, encode the spike events into a COO format sparse tensor, and directly trigger the sparse matrix multiplication of the Cube Unit.
[0121] In the above embodiments, the analog signals are collected through the signal processing module and converted into digital signals, including analog-to-digital conversion and anti-aliasing filtering, ensuring the quality of the signals during the conversion process. The preprocessing steps include noise reduction and enhancement, which help improve the accuracy and effect of digital signal processing. The network construction module is responsible for constructing a spiking neuron network according to the number of points of the Fourier transform. By receiving spike inputs, calculating weighted sums, updating the membrane potential, and triggering spikes, it can efficiently process signals and convert them into frequency domain assignments, which is crucial for analyzing the frequency components of signals; using sparsity analysis to identify the sparse regions of signals and neuron activities can reduce the computational amount, improve the processing speed, and maintain the accuracy of signal analysis at the same time. The model conversion module deploys the spiking neuron network onto a neural network processor. By converting the time-driven characteristics into a static computational graph and simulating the surrogate gradient of spike activation, it also includes weight quantization and synaptic pruning optimization, which helps improve the running efficiency and response speed of the model. In the deployment stage, using the dynamic sparse acceleration ability of the neural network processor, encoding the spike events into a COO format sparse tensor and directly triggering the sparse matrix multiplication of the Cube Unit can greatly improve the computational efficiency, especially when dealing with large-scale data.
[0122] Embodiment 10:
[0123] As shown Figure 6 in the figure, based on Embodiments 1-8, the application of the method for fast Fourier transform based on a pulsed neuron network provided by the embodiments of the present invention in a neural network processor includes:
[0124] A data acquisition module, which after parallelly acquiring analog signals, digital signals, and high-speed pulse signals by a multi-channel interface compatible circuit, directly connects to the NPU memory through PCIe 3.0 x4 or the Ascend dedicated AXI bus;
[0125] A preprocessing module, which is used to decompose the preprocessing tasks into multiple subtasks and perform parallel processing on the Ascend chip to improve the processing speed and achieve noise reduction and enhancement simultaneously;
[0126] A detection and optimization module, which is used to identify sparse regions of input data or neuron activities using a sparsity detection algorithm (such as threshold-based sparsity analysis);
[0127] A model construction module, which is used to construct an SNN (refer to the appendix Figure 7 ) for implementing the FFT algorithm, including a network structure design unit, a neuron model selection unit, and a connection weight setting unit.
[0128] A model conversion module, which is used to convert the SNN model into a format supported by the Ascend chip and utilize the inference engine of Ascend for acceleration;
[0129] A power consumption management module, which is used to dynamically adjust parameters such as the working frequency and clamping degree of the computing unit according to the real-time computing load to optimize the computing performance and power consumption;
[0130] Among them, the power consumption control module realizes full-stack energy efficiency optimization based on the heterogeneous perception architecture of the Ascend chip: by real-time monitoring the input signal sparsity of the data acquisition module, the computing load of the preprocessing module, and the neuron activation rate of the model construction module, a resource demand profile is dynamically constructed; when the detection and optimization module identifies a high-sparsity region, a hierarchical regulation strategy is triggered - at the chip level, the DVFS module of the AI Core dynamically adjusts the voltage and frequency according to the number of FFT points, and at the task level, the task scheduler migrates low-priority processes to the low-power CPU cluster, and at the same time enables the sparse calculation skip function of the model construction module, and only allocates computing resources to valid pulse events to achieve energy-saving effects. At the same time, the SNN network layer in the model construction is dynamically compressed through the spectral energy distribution to reduce memory occupancy.
[0131] In the above embodiments, the Ascend 310 chip (Huawei Ascend series), as a processor designed specifically for artificial intelligence and high-performance computing, has significant advantages in supporting the fusion calculation of spiking neural networks (SNN) and fast Fourier transform (FFT) based on spiking neural networks; it has characteristics such as high energy efficiency ratio and real-time guarantee of hardware-level support for spatio-temporal coding; as an artificial intelligence dedicated SoC (System-on-Chip) developed independently by Huawei for edge computing scenarios, the Ascend 310 is based on a 7nm manufacturing process. Under the constraint of a typical 8W thermal design power consumption, it can provide a peak computing power of up to 16 TOPS (INT8) and 8 TFLOPS (FP16), and its energy efficiency ratio index reaches 2 TOPS / W (INT8), which is significantly better than similar edge AI acceleration chips.
[0132] In this embodiment, the pulse-domain FFT mapping model decomposes the butterfly operation into an SNN synaptic weight matrix and a membrane potential dynamic equation, and uses pulse timing coding to replace the complex multiplier; the heterogeneous acceleration architecture of the Ascend chip realizes the hardware-level sparse calculation of SNN-FFT through expanding the sparse tensor instruction set, the event-driven routing module, and dynamic voltage and frequency scaling (DVFS); the multi-module collaborative system includes data acquisition (multi-channel signal PCIe / AXI direct connection), preprocessing (parallel noise reduction and enhancement), detection optimization (threshold sparse analysis), model inference (pulse-amplitude decoding), model conversion (quantization deployment), and power consumption control modules. This device supports the processing of 4096-point FFT at the 0.8ms level on the Ascend chip, with an energy efficiency ratio of 18.6 GOPS / W, which is more than 5 times higher than the GPU solution. It is suitable for low-power real-time scenarios such as brain-computer interfaces and radar signal processing, breaking through the "memory wall" and computing density bottlenecks of traditional architectures. It has the characteristics of high-efficiency computing, low latency, and low power consumption, and is suitable for application fields such as large-scale data processing, real-time signal processing, image processing, and communication systems, and can significantly improve the power consumption of FFT operations to meet the requirements of high-efficiency real-time computing.
[0133] Figure 8 The block diagram of an exemplary electronic device suitable for implementing the embodiments of the present invention is shown.
[0134] The electronic device may include a central processing unit / microprocessor / master control chip, etc. 4; a storage medium 5, coupled to the central processing unit / microprocessor / master control chip, etc. 4, and storing computer-executable instructions therein for performing the steps of the various methods of the embodiments of the present invention when executed by the processor.
[0135] The central processing unit / microprocessor / master control chip, etc. 4 may include, but are not limited to, for example, one or more processors or microprocessors, etc.
[0136] The storage medium 5 may include, but is not limited to, for example, random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, computer storage media (such as hard disks, floppy disks, solid-state drives, removable disks, CD-ROMs, DVD-ROMs, Blu-ray discs, etc.).
[0137] In addition, the electronic device may further include (but is not limited to) a data bus 6, an input / output bus / external bus / device bus, etc. 7, a display 8, and input / output devices 9 (such as a keyboard, a mouse, a speaker, etc.).
[0138] The central processing unit / microprocessor / master control chip, etc. 4 can communicate with external devices (8, 9, etc.) via the I / O bus 7 through a wired or wireless network (not shown).
[0139] The storage medium 5 can also store at least one computer-executable instruction for performing the various functions and / or method steps in the embodiments described in the present technology when run by the central processing unit / microprocessor / master control chip, etc. 4.
[0140] In one embodiment, the at least one computer-executable instruction can also be compiled into or form a software product, and when one or more computer-executable instructions are run by a processor, the various functions and / or method steps in the embodiments described in the present technology are performed.
[0141] Figure 9 A schematic diagram of a computer-readable storage medium according to an embodiment of the present invention is shown.
[0142] As Figure 9 shown, instructions are stored on the computer-readable storage medium 11, and the instructions are, for example, computer-readable instructions 10. When the computer-readable instructions 10 are run by a processor, the various methods described above can be executed. The computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may, for example, include random access memory (RAM) and / or cache memory, etc. Non-transitory non-volatile memory may, for example, include read-only memory (ROM), hard disks, flash memory, etc. For example, the computer-readable storage medium 11 can be connected to a computing device such as a computer, and then, when the computing device runs the computer-readable instructions 10 stored on the computer-readable storage medium 11, the various methods described above can be performed.
[0143] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0144] The units described as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0145] In addition, each functional unit in various embodiments of the present invention can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0146] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The foregoing storage medium includes: USB flash drives, mobile hard disks, read-only memories (English full name: Read-Only Memory, English abbreviation: ROM), random access memories (English full name: Random Access Memory, English abbreviation: RAM), magnetic disks or optical disks and other various media that can store program codes.
[0147] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of various embodiments of the present invention.
Claims
1. A method for fast Fourier transform based on a pulsed neuron network, characterized in that, It includes the following steps: Collect an analog signal at a configured sampling rate parameter through an analog-to-digital converter, convert the analog signal into a digital signal, and the converted digital signal enters a field-programmable gate array after passing through an anti-aliasing filter; Perform noise reduction and enhancement preprocessing on the digital signal entering the field-programmable gate array to obtain a preprocessed digital signal; Construct a spiking neural network according to the number of points of the fast Fourier transform based on the spiking neural network; Use sparsity analysis to identify the sparse regions of the preprocessed digital signal and neuron activity; Deploy the spiking neural network on a neural network processor, expand the time-driven characteristic of the spiking neural network into a static computational graph, and simulate the surrogate gradient of spiking activation through a custom operator; use the neural network processor offline model conversion tool for weight quantization and synaptic pruning optimization; Among them, the process of simulating the surrogate gradient of spiking activation through a custom operator includes the following steps: When the membrane potential of a neuron exceeds the threshold, use the dynamic scaling characteristic of the hyperbolic tangent function as the surrogate gradient kernel. The hyperbolic tangent function forms a continuously differentiable saturation region near the spiking trigger threshold, and its radius of curvature is negatively correlated with the refractory period characteristic of the neuron, so that the pseudo-derivative at the threshold crossing point can faithfully reflect the temporal firing pattern of biological neurons; Introduce a time expansion factor to map the time-domain discrete events of the spike train to virtual connection paths in a multi-dimensional tensor space; each spike event generates a virtual wire with an attenuation factor along the time axis, and the virtual wire forms a dynamically conductive differential channel in the surrogate gradient field; When a spiking neuron is triggered, discrete coordinates in COO format are generated in the forward propagation, and in the backpropagation stage, a surrogate gradient field under four-dimensional spatio-temporal constraints is reconstructed through the discrete coordinates. The surrogate gradient field forms a conjugate match with the partial differential of the spike density function; adopt threshold modulation similar to the photoelectric effect to make the amplitude of the surrogate gradient dynamically adjust with the inter-spike interval to match the asynchronous computing characteristics of the Cube Unit in the neural processor.
2. The method for fast Fourier transform based on a pulsed neuron network according to claim 1, wherein The process of obtaining the preprocessed digital signal includes the following steps: The multi-channel interface is connected to the memory of the neural network processor, and the analog signal is collected in parallel through the multi-channel interface; the digital signal entering the field-programmable gate array is decomposed into a noise reduction sub-task and an enhancement sub-task; The noise reduction sub-task decomposes the wavelet threshold denoising algorithm into multiple 3×3 convolution kernels and processes 8 channels of signals in parallel on the neural network processor carried by the wavelet threshold denoising algorithm; The enhancement sub-task uses the vector unit to perform normalization and dynamic range compression; Obtain the digital signal after the noise reduction sub-task and the enhancement sub-task and input it into the spiking neural network.
3. The method for fast Fourier transform based on a pulsed neuron network according to claim 2, wherein The processing process of the noise reduction sub-task and the enhancement sub-task includes the following steps: The input digital signal is segmented by time window, and each frame of digital signal is modulated by a complex adjustable basis function and decomposed into high-frequency and low-frequency components; according to the local transient energy gradient of the digital signal, the threshold intensity is dynamically adjusted to achieve discriminative suppression in an environment mixed with impulsive noise and steady-state noise; The wavelet band-pass filtering is equivalent to a 3×3 reconfigurable convolution kernel array, and the parameters inside the kernel are updated frame by frame according to the time-frequency ridge line characteristics of the signal. For continuous digital signals, a smoothing kernel is adopted; for transient pulse digital signals, it is switched to a differential kernel. The vector processing unit carried by the neural network processor adopts a mixed-precision serial computing stream. The input signal is block-filtered with zero-phase shift, the dynamic gain curve is calculated in the time domain, and the non-linear energy compression is realized by using the logarithmic Hilbert transform of the digital signal envelope.
4. The method for fast Fourier transform based on a pulsed neuron network according to claim 1, wherein, The process of performing noise reduction and enhancement preprocessing also includes the following steps: The input digital signal is regularly blocked into 16×16 to obtain data blocks; the data blocks are alternately stored in 4 independent memory banks by using memory bank interleaving; the SIMD parallel capability of the NEON instruction set is used, and a single VMLA instruction simultaneously completes 4 groups of complex multiplication operations, and data duplication loading is avoided through register reuse technology. The scalar unit of the neural network processor adopts sliding window statistics, calculates the mean / standard deviation of the signal amplitude within the window in real time, accelerates the standard deviation calculation by using the hardware square root unit, generates a dynamic threshold and synchronizes it through the broadcast bus. The tensor unit parallelly loads 16×16 data blocks, uses vector comparison instructions to perform threshold judgment, generates a sparse mask, executes BITPACK compression, and stores non-zero data and coordinates in CSC format. A scalar-tensor double-buffered pipeline is constructed. The threshold calculation of the current frame is parallel to the sparse processing of the previous frame, and data is automatically transferred through DMA. The threshold bus ensures calculation synchronization; the dynamic range of the signal is detected in real time, the FP16 calculation mode is switched for small dynamic ranges, the all-zero mask triggers the calculation to skip, and the voltage and frequency are dynamically adjusted according to the load.
5. The method for fast Fourier transform based on a pulsed neuron network according to claim 1, wherein The process of constructing a spiking neuron network includes the following steps: A spiking neural network is constructed according to the number of points of the fast Fourier transform based on the spiking neuron network. Each layer contains a pair of leaky integrate-and-fire neurons, and the pair of leaky integrate-and-fire neurons processes the real and imaginary parts of each group of complex numbers. The rotation factor is decomposed into a real part and an imaginary part, and the weight matrix of each neuron pair is defined; after the input spiking signal is integrated by the weight, the input current is calculated. The butterfly span is calculated, and the butterfly distance and the output of the butterfly node of a certain layer are calculated; the last layer is spiking frequency encoding, that is, the average number of spiking times of the output layer is counted, and the frequency domain amplitude is obtained after normalization, completing the fast Fourier transform based on the spiking neuron network.
6. The method of fast Fourier transform based on a pulsed neuron network according to claim 5, wherein The process in which a pair of leaky integrate-and-fire neurons processes the real and imaginary parts of each group of complex numbers includes the following steps: The spiking output signal of the leaky integrate-and-fire neuron flows through the decoding module. Through the differential receiver at the dendritic input end, the spiking cluster is split into signals and fed into the dendritic circuit array; the enable signal triggers the intelligent gate driver of the driving module to activate the time gate synchronization sequence. The real part data stream is injected into the capacitive network of the dendritic circuit, and the weight is adaptively adjusted through the charge domain effect. The imaginary part path completes the amplitude-phase correction inside the driving module. The cell body circuit executes, and the Schmitt trigger compares the RC integration values of two paths in real time. When the vector amplitude exceeds the threshold of the programmable comparator, it triggers the transition latch to update the rotation phase parameter and sends an encoded pulse to the output bus through the final-stage driver.
7. The method of fast Fourier transform based on a pulsed neuron network according to claim 1, wherein The process of using sparsity analysis to identify the sparse regions of the preprocessed digital signal and neuron activity includes the following steps: In the spiking neuron network, the digital signal is decomposed into sine waves of different frequencies through the fast Fourier transform based on the spiking neuron network to obtain the frequency-domain representation of the signal; in the spiking neuron network, the L1 norm of the spike signals of the input layer neurons is calculated to obtain the sparse metric value of the signal; Through the sparsity analysis of the spike signals of the input layer neurons, the important features of the digital signal are extracted as the sparse region; By performing threshold processing on the spike signals of the output layer neurons, the neurons with a sparsity higher than the threshold are found, and the corresponding spike signals of the neurons are the sparse regions of the digital signal.
8. A system for fast Fourier transform based on a spiking neural network, which is used to implement the method for fast Fourier transform based on a spiking neural network according to any one of claims 1-7, characterized in that, The system based on the fast Fourier transform of the spiking neuron network includes: A signal processing module for collecting an analog signal through an analog-to-digital converter at the configured sampling rate parameter, converting the analog signal into a digital signal, and the converted digital signal enters the field programmable gate array after passing through an anti-aliasing filter; The digital signal entering the field programmable gate array is preprocessed for noise reduction and enhancement to obtain the preprocessed digital signal; A network construction module for constructing a spiking neuron network according to the number of points of the fast Fourier transform based on the spiking neuron network. Each layer of neurons in the spiking neuron network receives the input of the layer circle pulses, updates the membrane potential and triggers spikes through weighted sum calculation, and the spike density of the output layer is converted into a frequency-domain assignment through integration; Using sparsity analysis to identify the sparse regions of the preprocessed digital signal and neuron activity; the spiking neuron network includes a network structure design unit, a neuron model selection unit, and a connection weight setting unit; A model conversion module for deploying the spiking neuron network on a neural network processor, expanding the time-driven characteristics of the spiking neuron network into a static computational graph, and simulating the surrogate gradient of spike activation through a custom operator; using the offline model conversion tool of the neural network processor for weight quantization and synaptic pruning optimization; when deploying, using the dynamic sparse acceleration ability of the neural network processor to encode the spike events into a COO format sparse tensor and directly trigger the sparse matrix multiplication of the Cube Unit.
9. The system for fast Fourier transform based on a pulsed neuron network according to claim 8, wherein The neural network processor is an AI chip.
Citation Information
Patent Citations
General coding method for spiking neurons based on electroencephalogram time-frequency characterization
CN117217267A
Information processing method, apparatus, electronic device, storage medium and program product
WO2023038414A1