Fast Fourier transform method and system based on spiking neural network
By constructing a fast Fourier transform based on pulsed neuron networks on a neural network processor, and using sparseness analysis and alternative gradient optimization calculation diagrams, the problems of waste of computing resources and inefficient hardware energy efficiency in traditional methods are solved, and efficient and low-power signal processing is achieved.
Patent Information
- Application Number
- CN202510448842.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-10
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-10
AI Technical Summary
Traditional fast Fourier transforms based on pulsed neuron networks have problems of wasted computing resources and inefficient hardware when processing sparse signals, especially in embedded devices, resulting in excessive power consumption.
Analog signals are collected and preprocessed by analog-to-digital converters, and fast Fourier transform based on pulsed neuron network is constructed. The sparse area of the signal is identified using sparseness analysis, and the pulsed neuron network is deployed on the neural network processor. The alternative gradient of pulse activation is simulated by custom operators, and the calculation graph is optimized and weighted and synaptic pruning is performed.
It significantly reduces the redundant computing load, improves the computing efficiency, reduces power consumption, and improves the accuracy and efficiency of signal processing.
Smart Images

Figure CN119939234A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of digital signal processing and artificial intelligence chip technology, and in particular to a method and system for fast Fourier transform based on a spiking neuron network. Background Art
[0002] Signal processing is a general term for the processing of various types of electrical signals according to various intended purposes and requirements. The processing of analog signals is called analog signal processing, and the processing of digital signals is called digital signal processing. In actual signal processing, the frequency domain energy distribution of most voice, bioelectric signals, mechanical vibration and other signals is highly sparse, that is, the effective information of the signal is concentrated in a few frequency bands, while the energy of the remaining frequency bands is negligible. Taking the voice signal as an example, its energy is mainly distributed in the narrow band corresponding to the fundamental frequency and harmonic components, while the energy of high-frequency noise and low-frequency environmental interference accounts for less than 10%. Similarly, in EEG signals, the characteristic high-frequency oscillations during epileptic seizures only occupy the local frequency band of 30Hz-100Hz. The commonly used technology for signal processing is the fast Fourier transform (FFT) based on the spiking neural network. The traditional fast Fourier transform algorithm based on the spiking neural network requires equal weight calculation of all signal frequency points in design, regardless of their actual energy contribution. This mode of indiscriminate calculation of the entire frequency band leads to waste of computing resources and low hardware energy efficiency. The waste of computing resources is mainly reflected in: for a signal of length N, the traditional fast Fourier transform based on the spiking neural network needs to execute Even if the signal has only k significant frequency bands, the computational complexity cannot be reduced. For example, when processing a 4096-point speech signal, the traditional fast Fourier transform based on the spiking neural network needs to complete 24,576 complex multiplications, while the actual effective frequency band accounts for only about 20%, which means that more than 19,000 multiplication operations are consumed on irrelevant frequency bands. The low energy efficiency of hardware is mainly reflected in: in embedded devices such as hearing aids and wearable sensors, redundant calculations directly increase power consumption. Actual measurements show that the energy consumption of the traditional fast Fourier transform based on the spiking neural network to process a 1-second speech signal on an ARM Cortex-M4 processor is 85mJ, of which more than 60% of the power consumption is used to calculate low-energy frequency bands.
[0003] Although the traditional fast Fourier transform based on the pulse neural network widely meets the frequency domain analysis, its processing method is not efficient for some application scenarios, especially when the key frequency points cannot be accurately captured, which easily causes unnecessary power consumption. Even if the useful information of the signal only exists in a small part of the data, in order to obtain the complete spectrum, all N data points still need to be calculated, which increases the computational burden; in particular, many signals in nature (speech, images) are sparse in the frequency domain, that is, only a few frequency components have significant energy. The fast Fourier transform based on the pulse neural network cannot automatically identify and utilize this sparsity, but treats all frequency bins equally, resulting in problems such as low computational efficiency, high latency, and large resource usage. Summary of the invention
[0004] In order to achieve the above object, the present invention adopts the following technical scheme: In one aspect of the present invention, a method for fast Fourier transform based on a spiking neural network is provided, comprising the following steps: The analog signal is collected by an analog-to-digital converter under the configured sampling rate parameters, and the analog signal is converted into a digital signal. The converted digital signal passes through an anti-aliasing filter and enters a field editable gate array; the digital signal entering the field editable gate array is pre-processed with noise reduction and enhancement to obtain a pre-processed digital signal; Constructing a spiking neural network based on the number of points of the fast Fourier transform of the spiking neural network; using sparsity analysis to identify sparse regions of preprocessed digital signals and neuronal activity; The spiking neuron network is deployed on the neural network processor, the time-driven characteristics of the spiking neuron network are expanded into a static computational graph, and the alternative gradient of the spiking activation is simulated through a custom operator; the weight quantization and synaptic pruning optimization are performed using the neural network processor offline model conversion tool.
[0005] In an optional implementation manner, the process of obtaining the preprocessed digital signal comprises the following steps: The multi-channel interface is connected to the memory of the neural network processor, and the analog signal is collected in parallel through the multi-channel interface; the digital signal entering the field editable gate array is decomposed into a noise reduction subtask and an enhancement subtask; The denoising subtask decomposes the wavelet threshold denoising algorithm into multiple 3×3 convolution kernels and processes 8 signals on the neural network processor equipped with the wavelet threshold denoising algorithm; the enhancement subtask uses the vector unit to perform normalization and dynamic range compression; The digital signals after the denoising subtask and the enhancement subtask are obtained and input into the spiking neuron network.
[0006] In an optional implementation, the processing of the denoising subtask and the enhancement subtask includes the following steps: The input digital signal is segmented according to the time window, and each frame of the digital signal is modulated by a complex adjustable basis function and decomposed into high-frequency and low-frequency components. According to the local transient energy gradient of the digital signal, the threshold strength is dynamically adjusted to achieve discriminative suppression in a mixed environment of pulse noise and steady-state noise. The wavelet bandpass filter is equivalent to a 3×3 reconfigurable convolution kernel array. The parameters in the kernel are updated frame by frame according to the time-frequency ridge characteristics of the signal. For continuous digital signals, a smoothing kernel is used; for transient pulse digital signals, a differential kernel is used. The vector processing unit of the neural network processor adopts a mixed-precision serial calculation flow, divides the input signal into blocks for zero-phase shift filtering, calculates the dynamic gain curve in the time domain, and uses the logarithmic Hilbert transform of the digital signal envelope to achieve nonlinear energy compression.
[0007] In an optional implementation manner, the process of performing noise reduction and enhancement preprocessing further includes the following steps: The input digital signal is divided into 16×16 regular blocks to obtain data blocks; the data blocks are alternately stored in 4 independent storage banks using memory bank interleaving; the SIMD parallel capability of the NEON instruction set is used to complete 4 sets of complex multiplication operations simultaneously with a single VMLA instruction, and register multiplexing technology is used to avoid repeated data loading; The scalar unit of the neural network processor uses sliding window statistics to calculate the mean / standard deviation of the signal amplitude in the window in real time, uses the hardware square root unit to accelerate the standard deviation calculation, generates dynamic thresholds and synchronizes through the broadcast bus; the tensor unit loads 16×16 data blocks in parallel, uses vector comparison instructions to make threshold judgments, generates sparse masks, performs BITPACK compression, and stores non-zero data and coordinates in CSC format; Build a scalar-tensor double-buffered pipeline. The current frame threshold calculation is parallel to the sparse processing of the previous frame. Automatic data transfer is achieved through DMA, and the threshold bus ensures calculation synchronization. Real-time detection of signal dynamic range, switching to FP16 calculation mode for small dynamic range, all-zero mask triggers calculation skipping, and dynamically adjusts voltage and frequency according to load.
[0008] In an optional implementation, the process of constructing a spiking neural network comprises the following steps: A spiking neural network is constructed based on the number of points of the fast Fourier transform based on the spiking neural network, each layer of which contains a leaky integral release neuron pair, and the leaky integral release neuron pair processes the real part and the imaginary part of the complex number for each group; Decompose the rotation factor into real and imaginary parts, and define the weight matrix of each neuron pair; calculate the input current after the input pulse signal is weighted and integrated; Butterfly span calculation, calculate the butterfly distance and butterfly node output of a certain layer; the last layer is pulse frequency encoding, that is, the average number of pulses in the output layer is counted, and the frequency domain amplitude is obtained after normalization to complete the fast Fourier transform based on the pulse neural network.
[0009] In an optional embodiment, the process of leaky integrate-and-release neuron pairs processing the real and imaginary parts of complex numbers for each group comprises the following steps: The pulse output signal of the leakage integration and release neuron flows through the decoding module, and the pulse cluster is split into signals and fed into the dendrite circuit array through the differential receiver at the dendrite input end; the enable signal triggers the intelligent gate driver of the drive module to activate the time gate synchronization sequence; The real data stream is injected into the capacitor network of the dendrite circuit, and the weights are adaptively adjusted through the charge domain effect; the imaginary path completes the amplitude and phase correction inside the driving module; The cell body circuit is executed, and the Schmitt trigger compares the RC integral values of the two paths in real time. When the vector amplitude exceeds the threshold of the programmable comparator, the transition latch is triggered to update the rotation phase parameter and send the encoding pulse to the output bus through the final driver.
[0010] In an optional embodiment, the process of using sparsity analysis to identify sparse regions of preprocessed digital signals and neuronal activities comprises the following steps: In the spiking neural network, the digital signal is decomposed into sine waves of different frequencies by fast Fourier transform based on the spiking neural network to obtain the frequency domain representation of the signal; in the spiking neural network, the sparse measurement value of the signal is obtained by calculating the L1 norm of the pulse signal of the input layer neuron; By performing sparsity analysis on the impulse signals of the input layer neurons, the important features of the digital signal are extracted as sparse regions; By performing threshold processing on the pulse signals of the output layer neurons, neurons with sparsity higher than the threshold are found. The pulse signals corresponding to the neurons are the sparse areas of the digital signal.
[0011] In an optional embodiment, the process of simulating the replacement gradient of pulse activation by a custom operator comprises the following steps: When the membrane potential of a neuron exceeds the threshold, the dynamic scaling characteristics of the hyperbolic tangent function are used as an alternative gradient kernel. The hyperbolic tangent function forms a continuously differentiable saturation region near the pulse triggering threshold, and its radius of curvature is negatively correlated with the refractory period characteristics of the neuron, so that the pseudo-derivative of the threshold crossing point can faithfully reflect the temporal discharge pattern of biological neurons. The time expansion factor is introduced to map the time-domain discrete events of the pulse sequence into virtual connection paths in the multidimensional tensor space; each pulse event generates a virtual wire with an attenuation factor along the time axis, and the virtual wire forms a dynamically conductive differential channel in the alternative gradient field; When a pulse neuron is triggered, discrete coordinates in the COO format are generated in the forward propagation. In the backward propagation stage, the alternative gradient field under the constraints of four-dimensional space-time is reconstructed through the discrete coordinates. The alternative gradient field forms a conjugate match with the partial differential of the pulse density function. The threshold modulation similar to the photoelectric effect is used to dynamically adjust the amplitude of the alternative gradient with the pulse interval to match the asynchronous computing characteristics of the Cube Unit in the neural processor.
[0012] Another aspect of the present invention provides a system for fast Fourier transform based on a spiking neural network, which is used to implement the method for fast Fourier transform based on a spiking neural network. The system for fast Fourier transform based on a spiking neural network comprises: The signal processing module is used to collect analog signals under the configured sampling rate parameters through an analog-to-digital converter, convert the analog signals into digital signals, and the converted digital signals enter the field editable gate array after passing through an anti-aliasing filter; the digital signals entering the field editable gate array are subjected to noise reduction and enhancement preprocessing to obtain preprocessed digital signals; The network construction module is used to construct a spiking neuron network according to the number of points of the fast Fourier transform based on the spiking neuron network. Each layer of neurons in the spiking neuron network receives the pulse input of the circle layer, updates the membrane potential and triggers the pulse through weighted sum calculation, and converts the pulse density of the output layer into the frequency domain assignment through integration; uses sparsity analysis to identify the sparse areas of the preprocessed digital signal and neuron activity; the spiking neuron network includes a network structure design unit, a neuron model selection unit and a connection weight setting unit; The model conversion module is used to deploy the spiking neural network on the neural network processor, expand the time-driven characteristics of the spiking neural network into a static calculation graph, and simulate the alternative gradient of pulse activation through a custom operator; use the neural network processor offline model conversion tool to perform weight quantization and synapse pruning optimization; use the dynamic sparse acceleration capability of the neural network processor during deployment, encode the pulse event into a COO format sparse tensor, and directly trigger the sparse matrix multiplication of the Cube Unit.
[0013] Another aspect of the present invention provides an application of the method of fast Fourier transform based on pulse neural network in a neural network processor, including: the neural network processor is an AI chip.
[0014] The present invention collects continuous time domain signals according to the set sampling rate parameters through an analog-to-digital converter (ADC); the output signal suppresses high-frequency noise through an anti-aliasing filter to ensure that the frequency band of subsequent processing is limited within an effective range; the digital signal is denoised and enhanced in a field programmable gate array (FPGA) to optimize the input quality of subsequent frequency domain analysis. SNN is constructed according to the number of FFT points, and the time domain signal is mapped to pulse density using its time encoding characteristics, and the frequency domain amplitude is output by integration; the redundant computing load is dynamically reduced by combining the spatiotemporal sparsity of the input signal and the neuron activation; the SNN level is determined (such as the input layer mapping time window, the hidden layer nonlinear transformation), and the (LIF) model is used to realize the dynamic update of the membrane potential and optimize the pulse transmission efficiency. The time-driven logic of the SNN is expanded into a static data flow graph, and the pulse non-differentiability problem is solved by replacing the gradient, which is compatible with the ONNX reasoning framework; the pulse event is encoded as a COO sparse tensor, which triggers the Cube Unit sparse matrix multiplication of the NPU (skipping the zero value calculation), and the throughput is significantly improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings: Figure 1 A flowchart of a method for fast Fourier transform based on a spiking neural network provided in Example 1 of the present invention; Figure 2 A process diagram of obtaining a preprocessed digital signal provided in Embodiment 2 of the present invention; Figure 3 A process diagram of constructing a spiking neural network provided in Example 3 of the present invention; Figure 4 A process diagram of simulating pulse activation substitution gradient by a custom operator provided in Example 4 of the present invention; Figure 5 A system block diagram of a fast Fourier transform based on a spiking neural network provided in Example 9 of the present invention; Figure 6 This is a schematic diagram of the application of the fast Fourier transform method based on the pulse neural network in the embodiment 10 of the present invention in the neural network processor; Figure 7 Schematic diagram of the LIF neuron circuit model in Example 10 of the present invention; Figure 8 A block diagram of an electronic device provided by the present invention; Fig. 9 A block diagram of a computer-readable storage medium provided for the present invention. DETAILED DESCRIPTION
[0016] The technical solutions in the embodiments of the present invention will be described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0017] In the following, the terms "first", "second", etc. are used only for convenience of description and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, a feature defined as "first", "second", etc. may explicitly or implicitly include one or more of the features. In the description of the present invention, unless otherwise specified, "plurality" means two or more.
[0018] In the present invention, unless otherwise clearly specified and limited, the term "connection" should be understood in a broad sense, for example, "connection" can be a fixed mechanical connection, or a detachable mechanical connection, or integrated; or, "connection" can be a direct connection, or an indirect connection through an intermediate medium. In addition, unless otherwise clearly specified and limited, the term "coupling" should be understood in a broad sense, for example, "coupling" can be a direct electrical connection, such as physical contact and electrical conduction between two components, and can also be understood as electrical connection between different components in a circuit structure through physical lines such as copper foil or wires on a printed circuit board (PCB) that can transmit electrical signals to transmit electrical signals; or, "coupling" can be an indirect electrical connection between two components through an intermediate medium; or, "coupling" can be an electrical connection between two components in an air-spaced / non-contact manner, for example, two components are electrically connected by capacitive coupling to transmit electrical signals.
[0019] In the embodiments of the present invention, directional terms such as "up", "down", "left" and "right" may be defined including but not limited to the orientation relative to the schematic placement of the components in the drawings. It should be understood that these directional terms may be relative concepts, which are used for relative description and clarification, and may change accordingly according to the change of the orientation of the components in the drawings.
[0020] Embodiment 1: like Figure 1 As shown, an embodiment of the present invention provides a method for fast Fourier transform based on a spiking neural network, comprising the following steps: Step S100: collecting analog signals at a configured sampling rate parameter through an analog-to-digital converter, converting the analog signals into digital signals, and passing the converted digital signals through an anti-aliasing filter and entering a field editable gate array; performing noise reduction and enhancement preprocessing on the digital signals entering the field editable gate array to obtain preprocessed digital signals; Step S200: constructing a spiking neuron network according to the number of points of the fast Fourier transform based on the spiking neuron network, wherein each layer of neurons in the spiking neuron network receives a circle layer pulse input, updates the membrane potential and triggers a pulse through weighted sum calculation, and converts the output layer pulse density into a frequency domain assignment through integration; using sparsity analysis to identify the sparse areas of the preprocessed digital signal and neuron activity; the spiking neuron network includes a network structure design unit, a neuron model selection unit and a connection weight setting unit; Step S300: deploy the spiking neuron network on the neural network processor, expand the time-driven characteristics of the spiking neuron network into a static computational graph, and simulate the alternative gradient of pulse activation through a custom operator; use the neural network processor offline model conversion tool to perform weight quantization and synapse pruning optimization; use the dynamic sparse acceleration capability of the neural network processor during deployment, encode the pulse event into a sparse tensor in COO format, and directly trigger the sparse matrix multiplication of the Cube Unit.
[0021] In the above embodiment, step S100 collects continuous time domain signals according to the set sampling rate parameters through an analog-to-digital converter (ADC); the output signal is passed through an anti-aliasing filter to suppress high-frequency noise to ensure that the frequency band of subsequent processing is limited within the effective range; the digital signal is denoised and enhanced in a field programmable gate array (FPGA) to optimize the input quality of subsequent frequency domain analysis. Step S200 constructs an SNN according to the number of FFT points, maps the time domain signal to pulse density using its time encoding characteristics, and outputs the frequency domain amplitude through integration; combines the spatiotemporal sparsity of the input signal and neuron activation to dynamically reduce redundant computing loads; determines the SNN level (such as input layer mapping time window, hidden layer nonlinear transformation), and uses the (LIF) model to dynamically update the membrane potential and optimize the pulse transmission efficiency. Step S300 expands the time-driven logic of the SNN into a static data flow graph, solves the pulse non-differentiability problem by replacing the gradient, and is compatible with the ONNX reasoning framework; encodes the pulse event into a COO sparse tensor, triggers the Cube Unit sparse matrix multiplication of the NPU (skips zero value calculation), and significantly improves the throughput.
[0022] Embodiment 2: like Figure 2 As shown, based on Example 1, the process of obtaining the preprocessed digital signal in step S100 provided in the embodiment of the present invention includes the following steps: Step S101: The multi-channel interface is connected to the memory of the neural network processor, and the analog signal is collected in parallel through the multi-channel interface; the digital signal entering the field editable gate array is decomposed into a noise reduction subtask and an enhancement subtask; Step S102: The denoising subtask decomposes the wavelet threshold denoising algorithm into multiple 3×3 convolution kernels, and processes 8 signals on the neural network processor equipped with the wavelet threshold denoising algorithm; the enhancement subtask uses the vector unit to perform normalization and dynamic range compression; Step S103: Obtain the digital signal after the denoising subtask and the enhancement subtask, and input it into the spiking neural network.
[0023] In the above embodiment, step S101 performs signal acquisition and subtask decomposition, and adopts a high-speed analog-to-digital converter and a time-division multiplexing mechanism to ensure synchronous sampling of multi-channel analog signals and reduce crosstalk and jitter; step S102 performs subtask parallel processing, with each core processing 2 signals and a total of 8 signals in parallel, thereby reducing memory access latency; vectorized execution of subtasks is enhanced, and single instructions are normalized to complete multi-data calculations; dynamic range compression reduces nonlinear distortion of high dynamic range data.
[0024] In this embodiment, ±10V analog signals are collected through a 16-channel differential analog-to-digital converter ADC, the sampling rate is configured to 1MSPS, and the signals are input to the FPGA after passing through an anti-aliasing filter. It also supports parallel data streams of the LVDS interface and is directly connected to the Ascend 310B development board through an FMC connector. For high-speed pulse sequences (such as neuromorphic camera output), they are input through a dedicated AXI-Stream interface, and timestamp compression encoding is used to reduce bandwidth occupancy. For data transmission, a PCIe 3.0 x4 interface or an Ascend chip-specific AXI bus is used to write the original data directly into the LPDDR4X memory pool of the Ascend chip, bypassing the CPU to reduce latency.
[0025] Embodiment 3: Based on Example 2, the processing process of the denoising subtask and the enhancement subtask in step S102 provided in the embodiment of the present invention includes the following steps: Step S1021: the input digital signal is segmented according to the time window, and each frame of the digital signal is modulated by a complex adjustable basis function to be decomposed into high-frequency and low-frequency components; according to the local transient energy gradient of the digital signal, the threshold strength is dynamically adjusted to achieve discriminative suppression in a mixed environment of impulsive noise and steady-state noise; Step S1022: The wavelet bandpass filter is equivalent to a 3×3 reconfigurable convolution kernel array, and the parameters in the kernel are updated frame by frame according to the time-frequency ridge characteristics of the signal. For continuous digital signals, a smoothing kernel is used; for transient pulse digital signals, a differential kernel is switched; Step S1023: The vector processing unit of the neural network processor adopts a mixed precision serial calculation flow, divides the input signal into blocks for zero-phase shift filtering, calculates the dynamic gain curve in the time domain, and uses the logarithmic Hilbert transform of the digital signal envelope to achieve nonlinear energy compression.
[0026] In the above embodiment, step S1021 complex adjustable basis function modulation breaks through the limitation of fixed wavelet basis, and realizes the optimal sparse representation of digital signal in the two-dimensional plane of time and frequency by dynamically adjusting the scale factor and phase offset of the basis function, thereby improving the frequency domain resolution and transient component capture capability; the transient energy gradient threshold dynamically generates a threshold curve based on the local microstructure characteristics of the digital signal, and suppresses the steady-state background noise under the premise of retaining the pulse signal. Step S1022 reconfigurable convolution kernel array realizes the real-time loading of filter kernel parameters through hardware programmable logic, so that a single set of convolution kernels has both smoothing and differential functions, avoiding the phase distortion problem of multi-stage filters in series in traditional methods; the parameter update guided by the time-frequency ridge feature uses the instantaneous frequency derivative of the signal as the basis for adjusting the kernel weight, so that the spatial domain convolution operation has the ability of joint time-frequency analysis. Step S1023 zero-phase shift filtering eliminates the group delay of traditional causal filtering through bidirectional recursive filtering, ensuring the time domain alignment of the signal, and is suitable for multi-sensor data fusion scenarios; logarithmic Hilbert transform compression maps the envelope energy of the digital signal to the logarithmic domain and performs analytical signal reconstruction, achieving equivalent dynamic range compression while avoiding frequency domain transformation.
[0027] Embodiment 4: On the basis of Example 1, the process of performing noise reduction and enhancement preprocessing in step S100 provided in the embodiment of the present invention further includes the following steps: Step S104: divide the input digital signal into 16×16 regular blocks to obtain data blocks; use memory bank interleaving to alternately store the data blocks into 4 independent storage banks; use the SIMD parallel capability of the NEON instruction set, a single VMLA instruction simultaneously completes 4 groups of complex multiplication operations, and avoids repeated data loading through register multiplexing technology, thereby improving the calculation speed of the rotation factor; Step S105: The scalar unit of the neural network processor uses sliding window statistics to calculate the mean / standard deviation of the signal amplitude in the window in real time, uses the hardware square root unit to accelerate the standard deviation calculation, generates a dynamic threshold and synchronizes through the broadcast bus; the tensor unit loads 16×16 data blocks in parallel, uses vector comparison instructions to make threshold judgments, generates sparse masks, performs BITPACK compression, and stores non-zero data and coordinates in CSC format; Step S106: Construct a scalar-tensor double buffer pipeline, the current frame threshold calculation is parallel to the previous frame sparse processing, data is automatically transferred through DMA, and the threshold bus ensures calculation synchronization; real-time detection of signal dynamic range, small dynamic range switching FP16 calculation mode, all-zero mask trigger calculation skipping, and dynamically adjust voltage frequency according to load.
[0028] In the above embodiment, step S104 is based on NEON's parallelized FFT preprocessing to solve the bottleneck of traditional FFT memory access and realize full pipeline of data supply and computing unit; it provides high-quality spectrum data aligned in time domain for sparse processing. Step S105 dynamically sparsely encodes and adaptively eliminates noise / redundant spectral components, compresses the data volume to 9.3%-22.1% of the original size, and saves 62% of memory bandwidth. Step S106 system-level dynamic optimization achieves a balance in scalar-tensor unit utilization, and reduces power consumption to 37%-52% in fixed mode to ensure real-time requirements.
[0029] Embodiment 5: like Figure 3 As shown, based on Example 1, the process of constructing a spiking neural network in step S200 provided in the embodiment of the present invention includes the following steps: Step S201: constructing a spiking neural network according to the number of points of the fast Fourier transform based on the spiking neural network, wherein each layer comprises a leakage integral release neuron pair, and the leakage integral release neuron pair processes the real part and the imaginary part of the complex number for each group; Step S202: decomposing the rotation factor into a real part and an imaginary part, defining a weight matrix for each neuron pair; calculating the input current after weighted integration of the input pulse signal; Step S203: butterfly span calculation, calculate the butterfly distance and butterfly node output of a certain layer; the last layer is pulse frequency coding, that is, the average number of pulses in the output layer is counted, and the frequency domain amplitude is obtained after normalization to complete the fast Fourier transform based on the pulse neural network.
[0030] Among them, according to the number of points N of the fast Fourier transform based on the pulse neural network, Layer SNN, the number of neurons in each layer is N / 2, corresponding to the number of butterfly operation nodes; for example, 8-point FFT requires a 3-layer structure, each layer contains 4 LIF neuron pairs (processing the real and imaginary parts of complex operations) LIF parameter settings: configure the membrane time constant of the LIF neuron =10ms, membrane potential threshold V th =1.0mV, the membrane potential update formula is:
[0031] in Indicates the current time The membrane potential, in millivolts (mV); Indicates the last moment The membrane potential; Indicates the current time The input current is in amperes (A); represents the membrane time constant, which indicates the response speed of the membrane potential to the change of input current, here it is 10 ms; According to the FFT rotation factor Set the synaptic weights; the weight matrix for each pair of neurons (real part, imaginary part) is split into: ,
[0032]
[0033] in Represents the complex exponential form, used for frequency domain analysis, k represents the frequency index, and N represents the total number of samples; represents the connection weights from real neurons to real neurons; Represents the connection weight from the real neuron to the imaginary neuron; Represents the connection weight from the imaginary neuron to the real neuron; represents the connection weight from the imaginary neuron to the imaginary neuron; Represents the complex plane rotation angle (unit: radians).
[0034] Each layer of neurons receives pulse input from the previous layer and calculates the weighted sum , updates the membrane potential and triggers a pulse. For example, The output of each butterfly node is: . Where Δ represents the butterfly span of the current layer; m represents the layer index of the current neural network; Indicates the current layer Butterfly operation nodes; represents the synaptic weight matrix; LIF(·) represents the leaky integrate-fire model; Represents the pulse stream received by the current neuron; represents the butterfly connection input across Δ neurons; Indicates the mth layer The output of the butterfly node is: the pulse density (number of pulses per unit time) of the final output layer is converted into the frequency domain amplitude through integration to complete the FFT calculation.
[0035] In the above embodiment, step S201 constructs a spiking neural network (SNN) based on the number of points of the fast Fourier transform (FFT) of the spiking neural network, and realizes parallel processing of complex signals by layered deployment of leaky integral release neuron pairs. Each group of neuron pairs processes the real part and imaginary part of the complex number respectively, ensuring the completeness of complex number operations in the pulse domain; it provides a hardware-based pulse computing infrastructure for the fast Fourier transform FFT based on the spiking neural network. Step S202 Decompose the rotation factor of the FFT into the corresponding even-symmetric component real part and the corresponding odd symmetric component imaginary part , and mapped to the weight matrix of the neuron pair; the input pulse signal is weightedly integrated to generate an input current as a pulse sequence to drive the membrane potential dynamics of the LIF neuron; the frequency domain operation is converted into a synaptic plasticity problem of pulse spatiotemporal coding. Step S203 calculates the span of the butterfly unit layer by layer, and realizes in-situ calculation by adjusting the pulse routing connection between neurons to reduce memory access overhead; the average firing rate of neurons is counted in the output layer, and the frequency domain amplitude is obtained after normalization; the nonlinear mapping of the time domain pulse sequence to the frequency domain amplitude is completed to meet the energy conservation characteristics of FFT.
[0036] In summary, this embodiment converts the complex multiplication and accumulation operation of FFT into the spatiotemporal event-driven calculation of the pulse neural network, which is suitable for low-power edge computing scenarios; the integration in the analog domain is realized through the leakage-threshold characteristics of the LIF neuron, which reduces the multiplication operations by about 90% compared with the digital FFT.
[0037] Embodiment 6: On the basis of Example 5, the process of leaky integral distribution neuron pairs processing the real part and imaginary part of a complex number for each group in step S201 provided in the embodiment of the present invention includes the following steps: Step S2011: The pulse output signal of the leakage integration and release neuron flows through the decoding module, and the pulse cluster is split into signals and fed into the dendrite circuit array through the differential receiver at the dendrite input end; the enable signal triggers the intelligent gate driver of the driving module to activate the time gate synchronization sequence; Step S2012: the real data stream is injected into the capacitor network of the dendrite circuit, and the weight is adaptively adjusted through the charge domain effect; the imaginary path completes the amplitude and phase correction inside the driving module; Step S2013: the cell body circuit is executed, and the Schmitt trigger compares the RC integral values of the two paths in real time. When the vector amplitude exceeds the threshold of the programmable comparator, the transition latch is triggered to update the rotation phase parameter, and the encoding pulse is sent to the output bus through the final driver.
[0038] In the above embodiment, step S2011 is pulse decoding and synchronous activation. The pulse output signal of the leakage integral release neuron is first processed by the decoding module; through the differential amplifier at the input end of the dendrite, the pulse signal is split into multiple sub-signals, which are fed into the dendrite circuit array respectively, realizing the spatial distribution of the pulse signal; at the same time, the enable signal triggers the intelligent gate driver in the driving module, activates the time gate synchronization sequence, ensures the timing control of signal processing, enables each sub-signal to be processed synchronously, and ensures the overall performance of the system. Step S2012 is real data injection and imaginary amplitude and phase correction. In the real data stream injection stage, the data signal is injected into the capacitor network of the dendrite circuit; the weight is adaptively adjusted through the charge domain effect, and the signal weight is dynamically adjusted to adapt to different input signal characteristics; the imaginary path completes the amplitude and phase correction inside the driving module. Amplitude and phase correction is to correct the amplitude and phase of the imaginary signal to eliminate distortion and delay in the signal transmission process and improve the accuracy of signal processing. Step S2013 is the cell body circuit execution and output pulse encoding, and the cell body circuit is responsible for the final processing of the execution signal. The Schmitt trigger compares the RC integral values of the real and imaginary paths in real time. When the vector amplitude exceeds the threshold of the programmable comparator, the transition latch is triggered to update the rotation phase parameter. The update of the rotation phase parameter realizes the dynamic adjustment of the signal vector, so that the signal vector always remains in the optimal state. Through the final driver, the encoded pulse signal is sent to the output bus to complete the signal processing and output.
[0039] Embodiment 7: Based on Example 1, the process of using sparsity analysis to identify sparse regions of preprocessed digital signals and neuronal activities in step S200 provided in the embodiment of the present invention includes the following steps: Step S204: in the spiking neural network, the digital signal is decomposed into sine waves of different frequencies by fast Fourier transform based on the spiking neural network to obtain a frequency domain representation of the signal; in the spiking neural network, the sparse measurement value of the signal is obtained by performing L1 norm calculation on the spiking signal of the input layer neuron; Step S205: extracting important features of the digital signal as a sparse region by performing sparsity analysis on the pulse signal of the input layer neuron; Step S206: By performing threshold processing on the pulse signals of the neurons in the output layer, neurons with sparsity higher than the threshold are found. The pulse signals corresponding to the neurons are the sparse regions of the digital signals.
[0040] In the above embodiment, step S204 is signal decomposition and sparsity measurement. In the pulse neural network, the digital signal in the time domain is decomposed into sine waves of different frequencies by fast Fourier transform based on the pulse neural network to obtain the frequency domain representation of the signal, and the digital signal is converted from the time domain to the frequency domain, which is convenient for analyzing the frequency characteristics and sparsity of the digital signal; the sparseness measurement value of the signal is obtained by calculating the L1 norm of the pulse signal of the input layer neuron; the purpose of the L1 norm calculation is to measure the sparsity degree of the digital signal, and the sparse signal will have a smaller value under the L1 norm; by calculating the L1 norm, the sparsity of the digital signal can be quantified. Step S205 is sparsity analysis and feature extraction. By performing sparsity analysis on the pulse signal of the input layer neuron, the important features of the digital signal are extracted, the key information and non-redundant parts in the digital signal are identified, and the basis for sparse area positioning is provided; sparsity analysis usually involves analyzing the energy distribution of the digital signal to find out which frequency components the energy is mainly concentrated on. The frequency components often correspond to the important features of the digital signal. By extracting these features, the digital signal is effectively sparsely encoded. Step S206: Threshold processing and sparse area positioning. By performing threshold processing on the pulse signals of the output layer neurons, neurons with sparsity higher than the threshold are found, and the most significant part of the digital signal, i.e., the sparse area, is identified; the pulse signals corresponding to neurons with sparsity higher than the threshold are the sparse areas of the digital signal. These sparse areas contain the core information of the signal and are of great significance to the characterization and processing of the signal; threshold processing is an effective sparsity quantification method. By setting a reasonable threshold, the sparse part of the digital signal can be distinguished from the non-sparse part, which facilitates the processing and analysis of digital signals.
[0041] In summary, this embodiment completes the entire process from signal decomposition, sparse measurement, feature extraction to sparse region positioning, providing technical support for identifying and processing sparse regions of digital signals; effectively extracting key information of the signal, reducing redundancy, and improving the efficiency and accuracy of signal processing.
[0042] Embodiment 8: like Figure 4 As shown, based on Example 1, the process of simulating the replacement gradient of pulse activation by a custom operator in step S300 provided in the embodiment of the present invention includes the following steps: Step S301: When the membrane potential of the neuron exceeds the threshold, the dynamic scaling characteristics of the hyperbolic tangent function are used as a replacement gradient kernel. The hyperbolic tangent function forms a continuously differentiable saturation region near the pulse triggering threshold, and its curvature radius is negatively correlated with the refractory period characteristics of the neuron, so that the pseudo-derivative of the threshold crossing point can faithfully reflect the temporal discharge pattern of the biological neuron; Step S302: introducing a time expansion factor to map the time-domain discrete events of the pulse sequence into virtual connection paths in the multidimensional tensor space; each pulse event generates a virtual wire with an attenuation factor along the time axis, and the virtual wire forms a dynamically conductive differential channel in the alternative gradient field; Step S303: When the pulse neuron is triggered, discrete coordinates in the COO format are generated in the forward propagation. In the backward propagation stage, the alternative gradient field under the four-dimensional space-time constraint is reconstructed by the discrete coordinates. The alternative gradient field forms a conjugate match with the partial differential of the pulse density function. The threshold modulation similar to the photoelectric effect is adopted to dynamically adjust the amplitude of the alternative gradient with the pulse interval to match the asynchronous computing characteristics of the Cube Unit in the neural processor.
[0043] In the above embodiment, step S301, the dynamic scaling characteristics of the hyperbolic tangent function transform the non-differentiable problem of discrete pulses into a continuously differentiable pseudo-gradient problem, ensuring the feasibility of back propagation; through the negative correlation between the radius of curvature and the refractory period, the inherent time characteristics of real neurons (such as repolarization delay) are simulated to make the gradient calculation conform to the law of discharge dynamics; a saturation zone is formed near the trigger threshold, which not only maintains the discreteness of the pulse event, but also retains the continuity of the gradient flow, solving the gradient disappearance problem of the traditional step function during training. Step S302 projects the discrete time domain pulse event into the multidimensional tensor space through the time expansion factor to form a computable virtual topological structure; the virtual wire with the attenuation factor simulates the time dependence of the biological synapse (such as short-term plasticity), ensuring that the gradient is dynamically adjusted with the pulse interval, and enhancing the ability to extract time series features; the conduction path of the replacement gradient field is updated instantly according to the pulse event, avoiding the redundant calculation of the traditional static calculation graph, and adapting to the asynchronous sparse pulse flow. Step S303 uses discrete coordinate sparse coding in COO format to significantly reduce the storage and computational overhead of back propagation, and is directly compatible with the sparse matrix multiplication unit of the brain-like chip; by reverse mapping the space-time constraints of discrete coordinates, the gradient propagation path is ensured to be strictly consistent with the forward propagation to avoid error diffusion; the pulse interval of the gradient amplitude is dynamically adjusted to simulate the adaptive change of synaptic weights with the discharge frequency, thereby realizing the coordinated optimization of hardware computing and biological models.
[0044] Embodiment 9: like Figure 5 As shown, on the basis of Embodiment 1 to Embodiment 8, the fast Fourier transform system based on the pulse neural network provided in the embodiment of the present invention comprises: The signal processing module 1 is used to collect analog signals under the configured sampling rate parameters through an analog-to-digital converter, convert the analog signals into digital signals, and the converted digital signals enter the field editable gate array after passing through an anti-aliasing filter; the digital signals entering the field editable gate array are subjected to noise reduction and enhancement preprocessing to obtain preprocessed digital signals; The network construction module 2 is used to construct a spiking neuron network according to the number of points of the fast Fourier transform based on the spiking neuron network. Each layer of neurons in the spiking neuron network receives the pulse input of the circle layer, updates the membrane potential and triggers the pulse through weighted sum calculation, and converts the pulse density of the output layer into the frequency domain assignment through integration; uses sparsity analysis to identify the sparse areas of the preprocessed digital signal and neuron activity; the spiking neuron network includes a network structure design unit, a neuron model selection unit and a connection weight setting unit; Model conversion module 3 is used to deploy the pulse neural network on the neural network processor, expand the time-driven characteristics of the pulse neural network into a static calculation graph, and simulate the alternative gradient of pulse activation through a custom operator; use the neural network processor offline model conversion tool to perform weight quantization and synapse pruning optimization; use the dynamic sparse acceleration capability of the neural network processor during deployment, encode the pulse event into a COO format sparse tensor, and directly trigger the sparse matrix multiplication of the Cube Unit.
[0045] In the above embodiment, the analog signal is collected by the signal processing module and converted into a digital signal, including analog-to-digital conversion and anti-aliasing filtering to ensure the quality of the signal during the conversion process. The preprocessing steps include noise reduction and enhancement, which help to improve the accuracy and effect of digital signal processing. The network construction module is responsible for constructing a pulse neuron network according to the number of points of the Fourier transform. By receiving pulse input, calculating weighted sums, updating membrane potentials, and triggering pulses, it can efficiently process signals and convert them into frequency domain assignments, which is crucial for analyzing the frequency components of signals; using sparsity analysis to identify sparse areas of signal and neuronal activity can reduce the amount of calculation, improve processing speed, and maintain the accuracy of signal analysis. The model conversion module deploys the pulse neuron network to the neural network processor, and by converting time-driven characteristics into static calculation graphs, simulating alternative gradients of pulse activation, and also including weight quantization and synaptic pruning optimization, it helps to improve the operating efficiency and response speed of the model. In the deployment stage, the dynamic sparse acceleration capability of the neural network processor is used to encode pulse events into COO format sparse tensors, directly triggering the sparse matrix multiplication of the Cube Unit, which can greatly improve the computing efficiency, especially when processing large-scale data.
[0046] Embodiment 10: like Figure 6 As shown, based on embodiments 1-8, the application of the fast Fourier transform method based on the pulse neural network provided in the embodiment of the present invention in the neural network processor includes: The data acquisition module uses a multi-channel interface compatible circuit to collect analog signals, digital signals, and high-speed pulse signals in parallel, and then directly connects to the NPU memory through the PCIe 3.0 x4 or Ascend dedicated AXI bus; The preprocessing module is used to decompose the preprocessing task into multiple subtasks and process them in parallel on the Ascend chip to improve the processing speed and achieve noise reduction and enhancement at the same time; A detection optimization module for identifying sparse regions of input data or neuronal activity using a sparsity detection algorithm (e.g., threshold-based sparsity analysis); Model building module, used to build SNN for implementing FFT algorithm (see Appendix Figure 7 ), including a network structure design unit, a neuron model selection unit and a connection weight setting unit.
[0047] Model conversion module, used to convert the SNN model into a format supported by Ascend chips and accelerate it using Ascend's inference engine; The power management module is used to dynamically adjust the operating frequency, fixture degree and other parameters of the computing unit according to the real-time computing load, so as to optimize the computing performance and power consumption; Among them, the power consumption control module realizes full-stack energy efficiency optimization based on the heterogeneous perception architecture of the Ascend chip: by real-time monitoring the input signal sparsity of the data acquisition module, the computing load of the preprocessing module and the neuron activation rate of the model construction module, a resource demand profile is dynamically constructed; when the detection optimization module identifies a high-sparse area, a hierarchical control strategy is triggered - at the chip level, the DVFS module of the AI Core dynamically adjusts the voltage frequency according to the number of FFT points, and at the task level, the task scheduler migrates low-priority processes to low-power CPU clusters, and enables the sparse calculation skipping function of the model construction module, which only allocates computing resources to valid pulse events to achieve energy-saving effects. At the same time, the SNN network layer in the model construction is dynamically compressed through the spectrum energy distribution to reduce memory usage.
[0048] In the above embodiment, the Ascend 310 chip (Huawei Ascend series), as a processor designed for artificial intelligence and high-performance computing, has significant advantages in supporting the fusion calculation of spiking neural network (SNN) and fast Fourier transform (FFT) based on spiking neural network; it has the characteristics of high energy efficiency and real-time hardware-level support for spatiotemporal coding; as Huawei's independently developed artificial intelligence dedicated SoC (System-on-Chip) for edge computing scenarios, Ascend 310 is based on a 7nm process technology. Under the typical 8W thermal design power consumption constraint, it can provide a peak computing power of up to 16TOPS (INT8) and 8TFLOPS (FP16), and its energy efficiency index reaches 2TOPS / W (INT8), which is significantly better than similar edge AI acceleration chips.
[0049] In this embodiment, the pulse domain FFT mapping model decomposes the butterfly operation into the SNN synaptic weight matrix and the membrane potential dynamic equation, and uses pulse timing coding to replace the complex multiplier; the Ascend chip heterogeneous acceleration architecture realizes the hardware-level sparse calculation of SNN-FFT by expanding the sparse tensor instruction set, event-driven routing module and dynamic voltage frequency adjustment (DVFS); the multi-module collaborative system includes data acquisition (multi-channel signal PCIe / AXI direct connection), preprocessing (parallel noise reduction enhancement), detection optimization (threshold sparse analysis), model reasoning (pulse-amplitude decoding), model conversion (quantized deployment) and power consumption control module. The device supports 0.8ms processing of 4096-point FFT on the Ascend chip, with an energy efficiency ratio of 18.6GOPS / W, which is more than 5 times higher than the GPU solution. It is suitable for low-power real-time scenarios such as brain-computer interface and radar signal processing, breaking through the "memory wall" and computing density bottleneck of traditional architecture. It has the characteristics of efficient computing, low latency and low power consumption. It is suitable for application fields such as large-scale data processing, real-time signal processing, image processing, communication systems, etc. It can significantly improve the power consumption of FFT operations and meet the needs of efficient real-time computing.
[0050] Figure 8 A block diagram of an exemplary electronic device suitable for implementing embodiments of the present invention is shown.
[0051] The electronic device may include a central processing unit / microprocessor / main control chip, etc. 4; a storage medium 5, coupled to the central processing unit / microprocessor / main control chip, etc. 4, and storing computer executable instructions therein, for performing the steps of each method of an embodiment of the present invention when executed by the processor.
[0052] The central processing unit / microprocessor / main control chip 4 may include but is not limited to, for example, one or more processors or microprocessors.
[0053] The storage medium 5 may include, but is not limited to, for example, random access memory (RAM), read-only memory (ROM), flash memory, EPROM memory, EEPROM memory, registers, computer storage media (such as hard disk, floppy disk, solid state drive, removable disk, CD-ROM, DVD-ROM, Blu-ray disc, etc.).
[0054] In addition, the electronic device may also include (but not limited to) a data bus 6, an input / output bus / external bus / device bus 7, a display 8, and input / output devices 9 (eg, keyboard, mouse, speaker, etc.).
[0055] The central processing unit / microprocessor / main control chip etc. 4 can communicate with external devices ( 8 , 9 etc.) through the I / O bus 7 via a wired or wireless network (not shown).
[0056] The storage medium 5 may also store at least one computer executable instruction for executing the various functions and / or method steps in the embodiments described in the present technology when executed by the central processing unit / microprocessor / main control chip 4.
[0057] In one embodiment, the at least one computer executable instruction may also be compiled into or constitute a software product, wherein one or more computer executable instructions are executed by a processor to perform the various functions and / or method steps in the embodiments described in the present technology.
[0058] Fig. 9 A schematic diagram of a computer-readable storage medium according to an embodiment of the present invention is shown.
[0059] like Fig. 9 As shown, instructions are stored on the computer-readable storage medium 11, and the instructions are, for example, computer-readable instructions 10. When the computer-readable instructions 10 are executed by the processor, the various methods described above can be executed. The computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. Non-temporary non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. For example, the computer-readable storage medium 11 can be connected to a computing device such as a computer, and then, when the computing device runs the computer-readable instructions 10 stored on the computer-readable storage medium 11, the various methods described above can be performed.
[0060] In the several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0061] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0062] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0063] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for executing all or part of the steps of the various embodiments of the present invention through a computer device (which can be a personal computer, server, or network device, etc.). The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (full name in English: Read-Only Memory, English abbreviation: ROM), random access memory (full name in English: Random Access Memory, English abbreviation: RAM), disk or optical disk, etc. Various media that can store program codes.
[0064] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for fast Fourier transform based on a spiking neural network, characterized in that: The following steps are involved: The analog signal is collected by the analog-to-digital converter under the configured sampling rate parameters, and the analog signal is converted into a digital signal. The converted digital signal passes through an anti-aliasing filter and enters the field editable gate array; The digital signal entering the field editable gate array is subjected to noise reduction and enhancement preprocessing to obtain a preprocessed digital signal; Constructing a spiking neural network based on the number of points of the fast Fourier transform of the spiking neural network; Use sparsity analysis to identify sparse regions of preprocessed digital signals and neuronal activity; The spiking neuron network is deployed on the neural network processor, the time-driven characteristics of the spiking neuron network are expanded into a static computational graph, and the alternative gradient of the spiking activation is simulated through a custom operator; the weight quantization and synaptic pruning optimization are performed using the neural network processor offline model conversion tool.
2. The method for fast Fourier transform based on spiking neural network as claimed in claim 1, characterized in that: The process of obtaining the preprocessed digital signal includes the following steps: The multi-channel interface is connected to the memory of the neural network processor, and the analog signal is collected in parallel through the multi-channel interface; the digital signal entering the field editable gate array is decomposed into a noise reduction subtask and an enhancement subtask; The denoising subtask decomposes the wavelet threshold denoising algorithm into multiple 3×3 convolution kernels and processes 8 signals on the neural network processor equipped with the wavelet threshold denoising algorithm; The enhancement subtask uses the vector unit to perform normalization and dynamic range compression; The digital signals after the denoising subtask and the enhancement subtask are obtained and input into the spiking neuron network.
3. The method for fast Fourier transform based on spiking neural network as claimed in claim 2, characterized in that: The processing of the denoising subtask and the enhancement subtask includes the following steps: The input digital signal is segmented according to the time window, and each frame of the digital signal is modulated by a complex adjustable basis function and decomposed into high-frequency and low-frequency components. According to the local transient energy gradient of the digital signal, the threshold strength is dynamically adjusted to achieve discriminative suppression in a mixed environment of pulse noise and steady-state noise. The wavelet bandpass filter is equivalent to a 3×3 reconfigurable convolution kernel array. The parameters in the kernel are updated frame by frame according to the time-frequency ridge characteristics of the signal. For continuous digital signals, a smoothing kernel is used; for transient pulse digital signals, a differential kernel is used. The vector processing unit of the neural network processor adopts a mixed-precision serial calculation flow, divides the input signal into blocks for zero-phase shift filtering, calculates the dynamic gain curve in the time domain, and uses the logarithmic Hilbert transform of the digital signal envelope to achieve nonlinear energy compression.
4. The method for fast Fourier transform based on spiking neural network as claimed in claim 1, characterized in that: The process of noise reduction and enhancement preprocessing also includes the following steps: The input digital signal is divided into 16×16 regular blocks to obtain data blocks; the data blocks are alternately stored in 4 independent storage banks using memory bank interleaving; the SIMD parallel capability of the NEON instruction set is used to complete 4 sets of complex multiplication operations simultaneously with a single VMLA instruction, and register multiplexing technology is used to avoid repeated data loading; The scalar unit of the neural network processor uses sliding window statistics to calculate the mean / standard deviation of the signal amplitude within the window in real time, uses the hardware square root unit to accelerate the standard deviation calculation, generates dynamic thresholds and synchronizes them through the broadcast bus; The tensor unit loads 16×16 data blocks in parallel, uses vector comparison instructions to make threshold judgments, generates sparse masks, performs BITPACK compression, and stores non-zero data and coordinates in CSC format; Build a scalar-tensor double-buffered pipeline. The current frame threshold calculation is parallel to the sparse processing of the previous frame. Automatic data transfer is achieved through DMA, and the threshold bus ensures calculation synchronization. Real-time detection of signal dynamic range, switching to FP16 calculation mode for small dynamic range, all-zero mask triggers calculation skipping, and dynamically adjusts voltage and frequency according to load.
5. The method for fast Fourier transform based on spiking neural network as claimed in claim 1, characterized in that: The process of building a spiking neural network includes the following steps: A spiking neural network is constructed based on the number of points of the fast Fourier transform based on the spiking neural network, each layer of which contains a leaky integral release neuron pair, and the leaky integral release neuron pair processes the real part and the imaginary part of the complex number for each group; Decompose the rotation factor into real and imaginary parts, and define the weight matrix of each neuron pair; calculate the input current after the input pulse signal is weighted and integrated; Butterfly span calculation, calculate the butterfly distance and butterfly node output of a certain layer; the last layer is pulse frequency encoding, that is, the average number of pulses in the output layer is counted, and the frequency domain amplitude is obtained after normalization to complete the fast Fourier transform based on the pulse neural network.
6. The method for fast Fourier transform based on spiking neural network as claimed in claim 5, characterized in that: The process of leaky integrate-and-release neuron pairs processing the real and imaginary parts of complex numbers for each group consists of the following steps: The pulse output signal of the leakage integration and release neuron flows through the decoding module, and the pulse cluster is split into signals and fed into the dendrite circuit array through the differential receiver at the dendrite input end; the enable signal triggers the intelligent gate driver of the drive module to activate the time gate synchronization sequence; The real data stream is injected into the capacitor network of the dendritic circuit, and the weights are adaptively adjusted through the charge domain effect; The imaginary part path completes the amplitude and phase correction inside the driver module; The cell body circuit is executed, and the Schmitt trigger compares the RC integral values of the two paths in real time. When the vector amplitude exceeds the threshold of the programmable comparator, the transition latch is triggered to update the rotation phase parameter and send the encoding pulse to the output bus through the final driver.
7. The method for fast Fourier transform based on spiking neural network as claimed in claim 1, characterized in that: The process of using sparsity analysis to identify sparse regions of preprocessed digital signals and neuronal activity includes the following steps: In the spiking neural network, the digital signal is decomposed into sine waves of different frequencies by fast Fourier transform based on the spiking neural network to obtain the frequency domain representation of the signal; in the spiking neural network, the sparse measurement value of the signal is obtained by calculating the L1 norm of the pulse signal of the input layer neuron; By performing sparsity analysis on the impulse signals of the input layer neurons, the important features of the digital signal are extracted as sparse regions; By performing threshold processing on the pulse signals of the output layer neurons, neurons with sparsity higher than the threshold are found. The pulse signals corresponding to the neurons are the sparse areas of the digital signal.
8. The method for fast Fourier transform based on spiking neural network as claimed in claim 1, characterized in that: The process of simulating the alternative gradient of pulse activation through a custom operator includes the following steps: When the membrane potential of a neuron exceeds the threshold, the dynamic scaling characteristics of the hyperbolic tangent function are used as an alternative gradient kernel. The hyperbolic tangent function forms a continuously differentiable saturation region near the pulse triggering threshold, and its radius of curvature is negatively correlated with the refractory period characteristics of the neuron, so that the pseudo-derivative of the threshold crossing point can faithfully reflect the temporal discharge pattern of biological neurons. The time expansion factor is introduced to map the time-domain discrete events of the pulse sequence into virtual connection paths in the multidimensional tensor space; each pulse event generates a virtual wire with an attenuation factor along the time axis, and the virtual wire forms a dynamically conductive differential channel in the alternative gradient field; When a pulse neuron is triggered, discrete coordinates in the COO format are generated in the forward propagation. In the backward propagation stage, the alternative gradient field under the constraints of four-dimensional space-time is reconstructed through the discrete coordinates. The alternative gradient field forms a conjugate match with the partial differential of the pulse density function. The threshold modulation similar to the photoelectric effect is used to dynamically adjust the amplitude of the alternative gradient with the pulse interval to match the asynchronous computing characteristics of the Cube Unit in the neural processor.
9. A system for fast Fourier transform based on a spiking neural network, used for implementing the method for fast Fourier transform based on a spiking neural network as claimed in any one of claims 1 to 8, characterized in that: The fast Fourier transform system based on the pulse neural network comprises: A signal processing module, used for collecting analog signals under configured sampling rate parameters through an analog-to-digital converter, converting the analog signals into digital signals, and the converted digital signals enter the field editable gate array after passing through an anti-aliasing filter; The digital signal entering the field editable gate array is subjected to noise reduction and enhancement preprocessing to obtain a preprocessed digital signal; The network construction module is used to construct a spiking neuron network according to the number of points based on the fast Fourier transform of the spiking neuron network. Each layer of neurons in the spiking neuron network receives the pulse input of the circle layer, updates the membrane potential and triggers the pulse through weighted sum calculation, and converts the pulse density of the output layer into frequency domain assignment through integration; Using sparsity analysis to identify sparse regions of preprocessed digital signals and neuron activities; the spiking neuron network includes a network structure design unit, a neuron model selection unit, and a connection weight setting unit; The model conversion module is used to deploy the spiking neural network on the neural network processor, expand the time-driven characteristics of the spiking neural network into a static calculation graph, and simulate the alternative gradient of pulse activation through a custom operator; use the neural network processor offline model conversion tool to perform weight quantization and synapse pruning optimization; use the dynamic sparse acceleration capability of the neural network processor during deployment, encode the pulse event into a COO format sparse tensor, and directly trigger the sparse matrix multiplication of the Cube Unit.
10. The fast Fourier transform system based on spiking neural network as claimed in claim 9, characterized in that: The neural network processor is an AI chip.
Citation Information
Patent Citations
General coding method for spiking neurons based on electroencephalogram time-frequency characterization
CN117217267A
Information processing method, apparatus, electronic device, storage medium and program product
WO2023038414A1
Cited By
Multi-source data health intervention method
CN120809211A
Low-power-consumption real-time signal processing method and system
CN120949918A
Calculation acceleration method based on cooperation of fast Fourier transform and neural network reasoning
CN120950266A
A fast fourier transform and neural network inference collaborative computing acceleration method
CN120950266B
Earthquake geomagnetic observation data acquisition and transmission method and system
CN121069473A