Neural mimicry signal driven video reconstruction method based on state-space equation

Through the video reconstruction method driven by neuromimicry signal based on state space equations, a multi-module video reconstruction network is built, which solves the problem of seamless connection between neuromimicry camera data and traditional computer vision technology, and achieves efficient and robust video reconstruction effect.

CN120147160AActive Publication Date: 2025-06-13UNIV OF SCI & TECH OF CHINA
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510215927.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-06-13
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently convert the asynchronous and sparse data of neuromimicry cameras into intensity images suitable for existing computer vision algorithms, making it difficult to seamlessly connect with traditional vision technologies.

Method used

Using a video reconstruction method driven by neuromimicry signal based on the state space equation, an efficient reconstruction of neuromimicry signals is achieved by constructing a video reconstruction network including shallow neuromimicry feature extraction module, spatial local enhancement downsampling module, spatial local enhancement upsampling module and video output module.

Benefits of technology

Efficient and robust video reconstruction is realized, the performance of neuromimicry data in model video reconstruction tasks is improved, and the shortcomings of the existing technology in space-time locality and translation invariance are solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147160A_ABST
    Figure CN120147160A_ABST
Patent Text Reader

Abstract

The invention discloses a neural mimicry signal driven video reconstruction method based on a state-space equation. The method comprises the following steps: 1, constructing a video reconstruction network based on the state-space equation; 2, introducing a random window displacement Mama module designed for the spatial domain characteristics of the neural mimicry signal; 3, introducing a Hilbert filling Mama module designed for the time-space domain characteristic of the neural mimicry signal; and 4, training the hybrid super-division network through back propagation, continuously optimizing until a loss function converges, and using an obtained video reconstruction model to reconstruct a to-be-processed neural mimicry signal so as to generate a corresponding high-quality video. Through the linear global modeling capability of the state-space equation and the targeted network module design, video reconstruction with efficient operation and excellent visual effect is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and particularly relates to reconstructing neuromorphic signals into video segments, and proposes a video reconstruction method driven by neuromorphic signals based on state space equations. Background Art

[0002] A neuromorphic camera is a novel vision sensor inspired by biological systems. Compared with traditional vision sensors, it has achieved significant technological breakthroughs in many aspects. The neuromorphic camera has advantages such as extremely high temporal resolution, excellent dynamic range, and ultra-low power consumption. Different from traditional cameras, the neuromorphic camera captures data in an asynchronous and sparse manner. Although this way can efficiently record the dynamic changes of the scene, it also makes its data difficult to directly interpret and difficult to be directly compatible with existing standard computer vision technologies. To overcome this technical obstacle, converting neuromorphic camera data into a more traditional intensity image becomes a key step, which helps to achieve seamless connection between the neuromorphic camera and traditional technologies in computer vision applications.

[0003] Currently, most mainstream computer vision methods rely on dense intensity frames captured by traditional CMOS sensors, which have limitations in temporal resolution and are difficult to fully capture fast-changing dynamic information. The emergence of the neuromorphic camera provides new possibilities for solving this problem, but its unique data form also poses new challenges to existing computer vision algorithms. Therefore, how to efficiently convert neuromorphic camera data into intensity images suitable for existing algorithms has become one of the important research directions. Summary of the Invention

[0004] The present invention is to solve the above-mentioned deficiencies of the existing technology, and proposes a video reconstruction method driven by neuromorphic signals based on state space equations, in order to achieve efficient operation, good reconstruction effect, and strong robustness of video reconstruction through the linear global modeling ability of state space equations and the design of targeted network modules.

[0005] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0006] A video reconstruction method driven by neuromorphic signals based on state space equations according to the present invention is characterized by including the following steps:

[0007] Step 1: Obtain the i-th neuromorphic signal stream with a time length of and the corresponding reference video frame , and perform voxelization operation on to obtain the i-th neuromorphic voxel , where and respectively represent the height and width of the neuromorphic voxel; C represents the number of channels after voxelization; represents the real number space;

[0008] Step 2: Construct a video reconstruction network based on the state space equation, including: a shallow neuromorphic feature extraction module, a spatial locality enhanced downsampling module, a spatio-temporal locality enhanced module, a spatial locality enhanced upsampling module, and a video output module; and process to obtain the i-th reconstructed video frame ;

[0009] Step 2.1: The shallow neuromorphic feature extraction module uses spatial convolutional residuals to extract features from the neuromorphic voxel to obtain the i-th shallow neuromorphic feature , where is the number of channels of the shallow neuromorphic feature;

[0010] Step 2.2: The spatial locality enhanced downsampling module consists of M cascaded random window downsampling modules, and after processing the shallow neuromorphic feature , it outputs the i-th set of neuromorphic features with enhanced spatial locality ; where represents the neuromorphic feature with enhanced spatial locality at the m-th level;

[0011] Step 2.3: The spatio-temporal locality enhanced module consists of N cascaded Mamba modules, and after processing the neuromorphic feature with enhanced spatial locality at the M-th level , it outputs the i-th set of neuromorphic features with enhanced spatio-temporal locality ; where represents the neuromorphic feature with enhanced spatio-temporal locality at the n-th level;

[0012] Step 2.4: The spatial locality enhanced upsampling module consists of M cascaded random window upsampling modules, and after processing the neuromorphic feature with enhanced spatio-temporal locality at the n-th level and , it outputs the i-th set of neuromorphic fusion features ; where represents the neuromorphic fusion feature at the m-th level;

[0013] Step 2.5: The video output module uses spatial convolutional residuals to process the neuromorphic fusion feature at the M-th level to obtain the i-th reconstructed video frame ;

[0014] Step 3: Construct a pre-trained fixed perception neural network and process and to obtain the i-th perception similarity metric ;

[0015] Step 4: Based on and and , construct the total loss function of the video reconstruction network based on the state space equation, and use the Adam optimizer to train the video reconstruction network based on the state space equation to update the network parameters until the total loss function converges, so as to obtain the optimal video reconstruction model for video reconstruction of the input neuromorphic signal.

[0016] Another feature of the video reconstruction method driven by neuromorphic signals based on the state space equation according to the present invention is that each cascaded random window downsampling module in step 2.2 includes: a random window displacement Mamba module, a downsampling module, and a convolutional long short-term memory network;

[0017] Step 2.2.1: When m = 1, the random window displacement Mamba module in the m-th cascaded random window downsampling module processes according to formula (1) to obtain the encoded feature after spatial local enhancement at the m-th level:

[0018] (1)

[0019] In formula (1), is the random window displacement layer, indicating that is randomly divided into a fixed number of patch blocks; is the multi-layer perceptron layer; is the visual Mamba module, and reshape is the reshaping operation; represents the random displacement amount, represents the random displacement amount in the height direction, represents the random displacement amount in the width direction; represents the uniform distribution, represents all possible displacements in the real number space, represents the spatial size of the random displacement window;

[0020] Step 2.2.2: When m = 1, the downsampling module in the m-th cascaded random window downsampling module processes to obtain the encoded feature after downsampling spatial enhancement at the m-th level;

[0021] Step 2.2.3: When m = 1, the convolutional long short-term memory network in the m-th cascaded random window downsampling module processes to obtain the neuromorphic features after the m-th level of spatial local enhancement ;

[0022] Step 2.2.4: When m = 2, 3, …, M, the neuromorphic features after the (m - 1)-th level of spatial local enhancement are input into the m-th cascaded random window downsampling module for processing, so as to obtain the neuromorphic features after the M-th level of spatial local enhancement .

[0023] Furthermore, each cascaded Mamba module in Step 2.3 includes: a random window displacement Mamba module and a Hilbert curve filling Mamba module;

[0024] Step 2.3.1: When n = 1, the n-th cascaded random window displacement Mamba module processes to obtain the bottleneck features after the n-th level of spatial local enhancement ;

[0025] Step 2.3.2: When n = 1, the n-th cascaded Hilbert curve filling Mamba module processes according to Equation (2) to obtain the bottleneck features after the n-th level of spatio-temporal local enhancement ;

[0026] (2)

[0027] In Equation (2), represents the Hilbert curve filling operation, indicating that is divided into a fixed number of patch blocks according to the filling method of the Hilbert curve; represents the position encoding added when processing the patch blocks; represents the operation of reshaping the patch blocks in reverse according to the Hilbert filling method;

[0028] Step 2.3.3: When n = 2, 3, ……, N, the bottleneck features after the (n - 1)-th level of spatio-temporal local enhancement are input into the n-th cascaded Mamba module for processing, so that the bottleneck features after the N-th level of spatial local enhancement are output by the N-th cascaded Mamba module.

[0029] Furthermore, each cascaded random window upsampling module in Step 2.4 includes: a random window displacement Mamba module and an upsampling module;

[0030] Step 2.4.1: When m = 1, after element-wise adding and , the neuromorphic features after the m-th level of local enhancement are obtained , and are input into the random window displacement Mamba module in the m-th cascaded random window upsampling module for processing to obtain the m-th level of spatio-temporal neuromorphic fusion features ;

[0031] Step 2.4.2: When m = 1, the upsampling module in the m-th cascaded random window upsampling module processes to obtain the m-th level of neuromorphic fusion features ;

[0032] Step 2.4.3: When m = 2, 3, …, M, after element-wise adding the (m - 1)-th level of neuromorphic fusion features and the (M - m + 1)-th level of spatially locally enhanced neuromorphic features , the neuromorphic features after the m-th level of local enhancement are obtained , and are input into the m-th cascaded random window upsampling module for processing, so that the M-th level of neuromorphic fusion features are output by the M-th cascaded random window upsampling module.

[0033] Furthermore, in the said Step 3, the i-th perceptual similarity index is obtained by using Equation (3):

[0034] (3)

[0035] In Equation (3), k represents the number of layers of the fixed perceptual neural network, K represents the total number of layers of the fixed perceptual neural network, represents the weight corresponding to the k-th layer of the fixed perceptual neural network, represents the feature vector extracted by the k-th layer of the fixed perceptual neural network for , represents the feature vector extracted by the k-th layer of the fixed perceptual neural network for , represents the i-th feature distance of the k-th layer of the fixed perceptual neural network.

[0036] Furthermore, in the said Step 4, the i-th loss function is established by using Equation (4), so as to obtain the total loss function , where I represents the total number of video frames;

[0037] (4)

[0038] In formula (4), is the temporal consistency loss, represents the number of the video frame when starting to calculate the temporal consistency loss, represents a hyperparameter, is the i-th perceptual loss function, and there is:

[0039] (5).

[0040] An electronic device according to the present invention includes a memory and a processor, characterized in that the memory is used to store a program for supporting the processor to execute the video reconstruction method, and the processor is configured to execute the program stored in the memory.

[0041] A computer-readable storage medium according to the present invention, characterized in that the computer program stored on the computer-readable storage medium executes the steps of the video reconstruction method when run by a processor.

[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0043] 1. Aiming at the problem that the existing visual state space model loses translational invariance under a fixed partition window, the present invention designs a random window offset strategy. By dynamically and randomly adjusting the window position, this strategy can better model the local information in the spatial domain, thereby improving the performance of neuromorphic data in the model video reconstruction task.

[0044] 2. To solve the deficiency of the existing visual state space model in spatio-temporal locality, the present invention proposes a Hilbert space filling curve mechanism specifically designed for the neuromorphic camera video reconstruction task. This mechanism can effectively capture the spatio-temporal relationship in the neuromorphic signal data, significantly enhancing the model's ability to model spatio-temporal information. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Figure 1 is a flowchart of a neuromorphic-driven video reconstruction method based on a state space equation in the present invention;

[0046] Figure 2a is a schematic diagram of a random window displacement Mamba in the present invention;

[0047] Figure 2b is a schematic diagram of a Hilbert filling Mamba in the present invention;

[0048] Figure 2c is a schematic diagram of a visual Mamba module in the present invention;

[0049] Figure 2d is a schematic diagram of a bidirectional visual Mamba module in the present invention;

[0050] Figure 3 It is a schematic diagram of random window displacement and Hilbert filling in the present invention;

[0051] Figure 4a It is a schematic diagram of the U-Net neural network of the overall network structure in the present invention;

[0052] Figure 4b It is the residual block network structure adopted by the shallow neuromorphic feature extraction module and the video output module in the present invention;

[0053] Figure 5 It is a performance result graph of the present invention compared with other different methods on a low-resolution dataset;

[0054] Figure 6 It is a performance result graph of the present invention compared with other different methods on a high-resolution dataset;

[0055] Figure 7 It is a performance result graph of the present invention compared with other different methods on multiple datasets;

[0056] Figure 8 Comparison graph of the computational resource consumption results of the video reconstruction method of the present invention on multiple datasets. Detailed implementation manner

[0057] In this embodiment, a video reconstruction method driven by neuromorphic signals based on the state space equation is based on the existing visual Mamba model. Combining the technical features of the video reconstruction task driven by neuromorphic signals, the structure of the visual Mamba model is optimized and the parameters are adjusted, so as to obtain an improved model that combines operation efficiency and reconstruction performance in the video reconstruction task, so that the reconstructed video is significantly improved in terms of temporal continuity and visual effect. Specifically, it includes the following steps: Figure 1 It gives the flowchart of the video reconstruction algorithm driven by neuromorphic signals based on the state space equation of the present invention.

[0058] Step 1: Obtain the i-th neuromorphic signal stream with a time length of and the corresponding reference video frame , and perform voxelization on to obtain the i-th neuromorphic voxel , where and respectively represent the height and width of the neuromorphic voxel; C represents the number of channels after voxelization; represents the real number space; at the same time, for the input neuromorphic voxel and the reference video frame Perform the same preprocessing process, including operations such as cropping, flipping, and rotating.

[0059] Step 2: Construct a video reconstruction network based on the state-space equation. The overall structure is as Figure 1 shown. The overall architecture of the network is a U-Net structure as shown in Figure 4a and includes: a shallow neuromorphic feature extraction module, a spatial locality enhanced downsampling module, a spatio-temporal locality enhanced module, a spatial locality enhanced upsampling module, and a video output module; and process to obtain the i-th reconstructed video frame ;

[0060] Step 2.1: The shallow neuromorphic feature extraction module uses the Figure 4b shown spatial convolutional residual to extract features from the neuromorphic voxels to obtain the i-th shallow neuromorphic feature , where is the number of channels of the shallow neuromorphic feature;

[0061] Step 2.2: The spatial locality enhanced downsampling module consists of M cascaded random window downsampling modules, and after processing the shallow neuromorphic feature , outputs the i-th set of neuromorphic features with enhanced spatial locality ; where represents the neuromorphic feature with enhanced spatial locality at the m-th level.

[0062] Step 2.2.1: When m = 1, the random window displacement Mamba module in the m-th cascaded random window downsampling module processes according to Equation (1) to obtain the encoded feature with enhanced spatial locality at the m-th level:

[0063] (1)

[0064] In Equation (1), is the random window displacement layer, indicating that is randomly divided into a fixed number of patch blocks; is the multi-layer perceptron layer; is the visual Mamba module, and reshape is the reshaping operation; represents the random displacement amount, represents the random displacement amount in the height direction, represents the random displacement amount in the width direction; represents the uniform distribution, represents all possible displacements in the real number space, represents the spatial dimension of the random displacement window; Figure 2a is a schematic diagram of the random window displacement Mamba, Figure 2c is a schematic diagram of the visual Mamba module VMB, Figure 2d is a schematic diagram of the bidirectional Mamba module, Figure 3 Part a in is a schematic diagram of the random window displacement strategy.

[0065] Step 2.2.2: When m = 1, the downsampling module in the m-th cascaded random window downsampling module processes to obtain the encoded feature after the m-th level of downsampling spatial enhancement ;

[0066] Step 2.2.3: When m = 1, the convolutional long short-term memory network in the m-th cascaded random window downsampling module processes to obtain the neuromorphic feature after the m-th level of spatial local enhancement ;

[0067] Step 2.2.4: When m = 2, 3,..., M, the neuromorphic feature after the (m - 1)-th level of spatial local enhancement is input into the m-th cascaded random window downsampling module for processing, so as to obtain the neuromorphic feature after the M-th level of spatial local enhancement.

[0068] Step 2.3: The spatio-temporal locality enhancement module is composed of N cascaded Mamba modules, and after processing the neuromorphic feature after the M-th level of spatial local enhancement, it outputs the i-th set of neuromorphic features after spatio-temporal locality enhancement ; where represents the neuromorphic feature after the n-th level of spatio-temporal locality enhancement;

[0069] Step 2.3.1: When n = 1, the n-th cascaded random window displacement Mamba module processes to obtain the bottleneck feature after the n-th level of spatial local enhancement ;

[0070] Step 2.3.2: When n = 1, the n-th cascaded Hilbert curve filling Mamba module processes according to Equation (2) to obtain the bottleneck feature after the n-th level of spatio-temporal locality enhancement ;

[0071] (2)

[0072] In Equation (2), Represents the Hilbert curve filling operation, indicating that is divided into a fixed number of patch blocks according to the filling method of the Hilbert curve; Represents the positional encoding added when processing the patch blocks; Represents the operation of reshaping the patch blocks in reverse according to the Hilbert filling method; Figure 2b Is a schematic diagram of the Hilbert curve filling Mamba, Figure 3 Part b in it is a schematic diagram of the Hilbert scanning curve.

[0073] Step 2.3.3: When n = 2, 3, ……, N, the bottleneck features after spatio-temporal local enhancement at the (n - 1)-th level Are input into the n-th cascaded Mamba module for processing, so that the bottleneck features after the N-th level of spatial local enhancement are output by the N-th cascaded Mamba module .

[0074] Step 2.4: The spatial locality enhancement upsampling module consists of M cascaded random window upsampling modules, and processes the neuromorphic features after the n-th level of spatio-temporal local enhancement and After processing, the i-th neuromorphic fusion feature set is output ; among them, Represents the m-th level of neuromorphic fusion features;

[0075] Step 2.4.1: When m = 1, and Are added element-wise to obtain the neuromorphic features after the m-th level of local enhancement , and are input into the random window displacement Mamba module in the m-th cascaded random window upsampling module for processing to obtain the m-th level of spatio-temporal neuromorphic fusion features ;

[0076] Step 2.4.2: When m = 1, the upsampling module in the m-th cascaded random window upsampling module processes To obtain the m-th level of neuromorphic fusion features ;

[0077] Step 2.4.3: When m = 2, 3, …, M, the m-th level of neuromorphic fusion features And the (M - m + 1)-th level of spatial local enhancement neuromorphic features Are added element-wise to obtain the neuromorphic features after the m-th level of local enhancement , and are input into the m-th cascaded random window upsampling module for processing, so that the M-th level of neuromorphic fusion features is output by the M-th cascaded random window upsampling module 。

[0078] Step 2.5: The video output module uses the spatial convolution residual as shown in Figure 4b shown in to process the M-level neuromorphic fusion features and obtain the i-th reconstructed video frame ;

[0079] Step 3: Construct a perceptual similarity calculation module, including a pre-trained fixed perceptual neural network, and process and and obtain the i-th perceptual similarity index according to Equation (3) ;

[0080] (3)

[0081] In Equation (3), k represents the number of layers of the fixed perceptual neural network, K represents the total number of layers of the fixed perceptual neural network, represents the weight corresponding to the k-th layer of the fixed perceptual neural network, represents the k-th layer of the fixed perceptual neural network for the extracted feature vector, represents the feature vector extracted by the k-th layer of the fixed perceptual neural network in the loss function for , represents the i-th feature distance of the k-th layer of the fixed perceptual neural network.

[0082] Step 4: Construct the total loss function of the video reconstruction network based on the state space equation , and use the Adam optimizer to train the video reconstruction network based on the state space equation to update the network parameters until the total loss function converges, so as to obtain the optimal video reconstruction model for video reconstruction of the input neuromorphic signal. Use Equation (4) to establish the i-th loss function , so as to obtain the total loss function , where I represents the total number of video frames;

[0083] (4)

[0084] In Equation (4), is the temporal consistency loss, represents the number of the video frame when calculating the temporal consistency loss starts, represents the hyperparameter, is the i-th perceptual loss function, and there is:

[0085] (5)

[0086] In this embodiment, an electronic device includes a memory and a processor. The memory is used to store a program that supports the processor to execute the above method, and the processor is configured to execute the program stored in the memory.

[0087] In this embodiment, a computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it executes the steps of the above method.

[0088] To quantitatively evaluate the effect of the present invention and verify its effectiveness, the method of the present invention was compared with a variety of video reconstruction algorithms driven by neuromorphic signals based on deep learning, including Convolutional Event Reconstruction Network (E2VID), Fast Event Reconstruction Network (FireNet), Fast Event Reconstruction Network+ (FireNet+), Convolutional Event Reconstruction Network+ (E2VID+), Spatial Adaptive Event Reconstruction Network (SPADE-E2VID), Sparse Event Reconstruction Network (Sparse-E2VID), and Transformer Event Reconstruction Network (ET-Net) and other methods. The experiments were carried out on the real shooting data of 3 public datasets, namely High-Quality Event Dataset (HQF), Robot Dataset (IJRR), and Mobile Video Sequences (MVSEC), aiming to comprehensively evaluate the performance of each method. To objectively measure the effect of the method, the following three performance metrics were selected as evaluation criteria: Mean Squared Error (MSE): It measures the pixel-level error between the reconstructed image and the real image. The smaller the MSE value, the higher the reconstruction quality; Structural Similarity Index (SSIM): It evaluates the similarity of the structural features of two images. The closer the SSIM is to 1, the better the structural consistency; Learned Perceptual Image Patch Similarity (LPIPS): It quantifies the visual perception difference. The lower the LPIPS value, the better the perceptual quality.

[0089] The numerical index evaluation is divided into three parts:

[0090] The first part is the image quality evaluation based on error sensitivity, and the evaluation criterion is the mean squared error MSE:

[0091] (5)

[0092] In Equation (5), represents the total number of pixels in the image, represents the j-th pixel value of the output video frame, Represents the j-th pixel value of the reference video frame.

[0093] The second part is based on the structural similarity of image quality assessment, and the evaluation criterion is the structural similarity index SSIM:

[0094] (6)

[0095] In Equation (6), is the mean value of is the mean value of is the variance of is the variance of is and the covariance of and are constants.

[0096] The third part is the learning perceptual image similarity assessment, and the evaluation index is LPIPS, as shown in Equation (3).

[0097] Figure 5 Shows the quantitative comparison results between the video reconstruction network of the present invention and existing methods. In terms of the MSE and SSIM metrics, the video reconstruction network proposed by the present invention is superior to all existing reconstruction networks based on neuromorphic signals. In terms of the LPIPS metric, the method of the present invention is also superior to most existing methods. Except for the high-quality event dataset, the LPIPS metric shows a slight decrease compared to the convolutional event reconstruction network + method. Overall, the method of the present invention performs excellently in terms of the MSE, SSIM, and LPIPS metrics and reaches the current state-of-the-art level on multiple datasets.

[0098] Figures 6 - 7 Shows the visual comparison effect of the present invention on the test set. Among them Figure 6Shows the visual comparison of the reconstruction results of the method proposed in the present invention with all baseline methods on the video segments of the low-resolution dataset, high-quality event dataset, robot dataset, and driving record dataset. In addition, the present invention also provides real image (GroundTruth, GT) frames as a reference. It can be observed that the reconstructed frames of the Fast Event Reconstruction Network+ and the Convolutional Event Reconstruction Network+ lack accuracy in brightness, resulting in poor overall visual quality. The reconstruction results of the Transformer Event Reconstruction Network are visually more realistic than those of the Fast Event Reconstruction Network+ and the Convolutional Event Reconstruction Network+, presenting an effect closer to the real scene. In contrast, the method proposed in the present invention can further enrich the final reconstruction results, showing more fine details, while significantly reducing the artifacts common in the Convolutional Event Reconstruction Network+. In addition, the image contrast of the reconstructed frames of the present invention is very close to that of the real image (GT). To further support the comparative analysis of the present invention, a series of high-definition video sequences were captured using the Prophesee EVK4 camera in this embodiment. Figure 7 Shows a series of comparison images. The method proposed in the present invention provides better visual effects and significantly fewer artifacts compared to other methods.

[0099] To verify the efficiency of the network of the present invention, the computational complexity of the present invention is analyzed compared with other methods in this embodiment, focusing on the total number of parameters and the inference time on the 3090 GPU. To make the comparison more meaningful, the present invention selects two resolution types of currently commonly used neuromorphic camera sensors: low resolution (L) and high resolution (H). The experimental results are as Figure 8 shown. In the table, the number of parameters of the model is in millions (M), and the inference time is expressed in milliseconds (ms). As Figure 8 can be seen, the method of the present invention achieves a good balance between computational complexity and performance. On the one hand, compared with the Transformer Event Reconstruction Network, the method of the present invention not only reduces the number of parameters by half, but also significantly reduces the inference time when processing large-sized inputs. On the other hand, compared with the CNN-based method, the method of the present invention achieves a significant performance improvement with only a slight increase in computational cost, fully demonstrating the superiority of the method of the present invention.

Claims

1. A video reconstruction method driven by a neuromorphic signal based on a state-space equation, characterized in that: The following steps are involved: Step 1: Get the time length The i-th neuromorphic signal flow And the corresponding reference video frame , and Perform voxelization to obtain the i-th neuromorphic voxel ,in, and They represent the height and width of the neuromorphic voxel respectively; C represents the number of channels after voxelization; represents the real number space; Step 2: Construct a video reconstruction network based on state-space equations, including: shallow neuromorphic feature extraction module, spatial locality enhancement downsampling module, spatiotemporal locality enhancement module, spatial locality enhancement upsampling module and video output module; and Processing is performed to obtain the i-th reconstructed video frame ; Step 2.1: The shallow neuromorphic feature extraction module uses spatial convolution residuals to extract neuromorphic voxels Perform feature extraction to obtain the i-th shallow neural mimicry feature ,in, is the number of channels of the shallow neuromorphic features; Step 2.2: The spatial locality enhanced downsampling module consists of M cascaded random window downsampling modules, and performs a deep neural mimicry on the shallow neural mimicry features. After processing, the i-th spatially locally enhanced neuromorphic feature set is output ;in, represents the neuromorphic features after local enhancement in the m-th level space; Step 2.3: The spatiotemporal locality enhancement module consists of N cascaded Mamba modules, and enhances the neuromorphic features of the M-th level spatial locality After processing, the i-th neuromorphic feature set with enhanced spatiotemporal locality is output ;in, represents the neuromorphic features after the nth level of spatiotemporal local enhancement; Step 2.4: The spatial locality enhanced upsampling module is composed of M cascaded random window upsampling modules, and the neuromorphic features after the n-th level of spatiotemporal local enhancement are and After processing, the i-th neuromorphic fusion feature set is output ;in, represents the m-th level neuromorphic fusion feature; Step 2.5: The video output module uses the spatial convolution residual to fusion features of the M-th level neuromorphic After processing, the i-th reconstructed video frame is obtained ; Step 3: Build a pre-trained fixed perceptron neural network and and Processing is performed to obtain the i-th perceptual similarity index ; Step 4: Based on and as well as , construct the total loss function of the video reconstruction network based on the state space equation , and use the Adam optimizer to train the video reconstruction network based on the state-space equation to update the network parameters until the total loss function Until convergence, the optimal video reconstruction model is obtained, which is used to reconstruct the video of the input neuromorphic signal.

2. The video reconstruction method driven by a neuromorphic signal based on a state-space equation according to claim 1, characterized in that: Each cascaded random window downsampling module in step 2.2 includes: a random window displacement Mamba module, a downsampling module, and a convolutional long short-term memory network; Step 2.2.1: When m=1, the random window shift Mamba module in the random window downsampling module of the mth cascade is converted according to formula (1) Processing is performed to obtain the encoding features after local enhancement of the mth level space : (1) In formula (1), is a random window displacement layer, which means Randomly divide into a fixed number of patches; It is a multi-layer perceptron layer; is the visual Mamba module, and reshape is the reshaping operation; represents the random displacement, represents the random displacement in the height direction, Represents the random displacement in the width direction; represents uniform distribution, represents all possible displacements in the real number space, represents the spatial size of the random displacement window; Step 2.2.2: When m=1, the downsampling module in the mth cascaded random window downsampling module is Processing is performed to obtain the encoding features after the mth level of downsampling space enhancement ; Step 2.2.3: When m=1, the convolutional long short-term memory network in the mth cascaded random window downsampling module is After processing, the neuromorphic features after local enhancement of the m-th level space are obtained. ; Step 2.2.4: When m=2,3,…,M, the neuromorphic features after local enhancement of the m-1th level space are The input is processed in the random window downsampling module of the mth cascade to obtain the neuromorphic features after local enhancement of the Mth level space. .

3. The video reconstruction method driven by a neuromorphic signal based on a state-space equation according to claim 2, characterized in that: Each cascaded Mamba module in step 2.3 includes: a random window displacement Mamba module and a Hilbert curve filling Mamba module; Step 2.3.1: When n=1, the random window shift Mamba module of the nth cascade Processing is performed to obtain the bottleneck features after local enhancement of the nth level space ; Step 2.3.2: When n=1, the Hilbert curve of the nth cascade fills the Mamba module according to formula (2) Processing is performed to obtain the bottleneck features after the nth level of spatiotemporal local enhancement ; (2) In formula (2), Represents the Hilbert curve filling operation, which means Divide into a fixed number of patches according to the filling method of the Hilbert curve; Indicates the position code added when processing the patch block; Represents the operation of reshaping the patch block in reverse according to the Hilbert filling method; Step 2.3.3: When n=2,3,…,N, the bottleneck features after the n-1th level of spatiotemporal local enhancement are The input is processed in the nth cascaded Mamba module, so that the Nth cascaded Mamba module outputs the bottleneck features after the Nth level of spatial local enhancement. .

4. The method for video reconstruction driven by a neuromorphic signal based on a state-space equation according to claim 3, characterized in that: Each cascaded random window upsampling module in step 2.4 includes: a random window shift Mamba module and an upsampling module; Step 2.4.1: When m=1, and After element-by-element addition, the neuromorphic features after local enhancement at the mth level are obtained. , and input the random window displacement Mamba module in the mth cascade random window upsampling module for processing to obtain the mth level spatiotemporal neuromorphic fusion feature ; Step 2.4.2: When m=1, the upsampling module in the mth cascaded random window upsampling module is Processing is performed to obtain the mth level of neuromorphic fusion features ; Step 2.4.3: When m=2,3,…,M, the m-1th level of neuromorphic fusion features and the M-m+1th level spatially local enhanced neuromorphic features After element-by-element addition, the neuromorphic features after local enhancement at the mth level are obtained. , and is input into the mth cascade random window upsampling module for processing, so that the Mth cascade random window upsampling module outputs the Mth level neuromorphic fusion feature .

5. The method for video reconstruction driven by a neuromorphic signal based on a state-space equation according to claim 4, characterized in that: In step 3, formula (3) is used to obtain the i-th perceptual similarity index: : (3) In formula (3), k represents the number of layers of the fixed perceptual neural network, K represents the total number of layers of the fixed perceptual neural network, represents the weight corresponding to the k-th layer of fixed perceptual neural network, represents the k-th layer fixed perception neural network pair The extracted feature vectors are Represents the k-th layer of the fixed perceptual neural network The extracted feature vectors are represents the i-th feature distance of the k-th layer fixed perceptron neural network.

6. The method for video reconstruction driven by a neuromorphic signal based on a state-space equation according to claim 5, characterized in that: In step 4, the i-th loss function is established using formula (4): , thus obtaining the total loss function , where I represents the total number of video frames; (4) In formula (4), is the temporal consistency loss, Indicates the number of the video frame from which the temporal consistency loss is calculated. represents the hyperparameter, is the i-th perceptual loss function, and we have: (5)。 7. An electronic device, comprising a memory and a processor, characterized in that: The memory is used to store a program that supports a processor to execute the video reconstruction method according to any one of claims 1 to 6, and the processor is configured to execute the program stored in the memory.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the video reconstruction method according to any one of claims 1 to 6 are executed.

Citation Information

Patent Citations

  • Butterfly image recognition method based on convolutional neural network

    CN115631417A

  • Node scale adaptive neural mimicry calculation neuron classification method

    CN118626903A

  • Contextual visual-based SAR target detection method and apparatus, and storage medium

    US20230184927A1

  • Image processing method, apparatus, and non-transitory computer-readable medium

    US20230325974A1