A video reconstruction method based on state-space equations and driven by neuromorphic signals.
By employing a state-space equation-based neuromorphic signal-driven video reconstruction method, and utilizing random window offset and Hilbert space-filling curve mechanisms, the compatibility problem between neuromorphic camera data and traditional computer vision techniques is solved, achieving efficient and robust video reconstruction results.
Patent Information
- Application Number
- CN202510215927.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-11-14
- Estimated Expiration
- 2045-02-26
AI Technical Summary
Existing computer vision methods struggle to effectively handle the asynchronous and sparse data from neuromorphic cameras, making it difficult to directly interpret the data and ensuring compatibility with traditional computer vision technologies. This hinders the seamless integration of neuromorphic cameras into computer vision applications.
A video reconstruction method based on state-space equations and driven by neuromorphic signals is adopted. By constructing a shallow neuromorphic feature extraction module, a spatial locality enhancement downsampling module, a spatiotemporal locality enhancement module, a spatial locality enhancement upsampling module, and a video output module, and combining a random window offset strategy and a Hilbert space-filling curve mechanism, efficient video reconstruction is achieved.
It significantly improves the performance of neuromorphic data in model video reconstruction tasks, enhances the ability to model spatiotemporal information, and improves the robustness and efficiency of reconstruction results.
Smart Images

Figure CN120147160B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of image processing technology, and in particular relates to the reconstruction of video segments from neuromorphic signals. It proposes a video reconstruction method driven by neuromorphic signals based on state-space equations. Background Technology
[0002] Neuromorphic cameras, a novel type of visual sensor inspired by biological systems, have achieved significant technological breakthroughs compared to traditional visual sensors in several aspects. Neuromorphic cameras possess advantages such as extremely high temporal resolution, excellent dynamic range, and ultra-low power consumption. Unlike traditional cameras, neuromorphic cameras capture data asynchronously and sparsely. While this approach efficiently records dynamic changes in a scene, it also makes the data difficult to interpret directly and incompatible with existing standard computer vision technologies. To overcome this technical obstacle, converting neuromorphic camera data into more traditional intensity images is a crucial step, which facilitates seamless integration of neuromorphic cameras with traditional technologies in computer vision applications.
[0003] Currently, most mainstream computer vision methods rely on dense intensity frames captured by traditional CMOS sensors. These sensors have limitations in temporal resolution, making it difficult to fully capture rapidly changing dynamic information. The emergence of neuromorphic cameras offers new possibilities for solving this problem, but their unique data format also presents new challenges to existing computer vision algorithms. Therefore, how to efficiently convert neuromorphic camera data into intensity images suitable for existing algorithms has become one of the important research directions. Summary of the Invention
[0004] This invention aims to address the shortcomings of existing technologies by proposing a neuromorphic-driven video reconstruction method based on state-space equations. The method aims to achieve efficient, effective, and robust video reconstruction through the linear global modeling capability of state-space equations and targeted network module design.
[0005] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:
[0006] The present invention provides a video reconstruction method driven by neuromorphic signals based on state-space equations, characterized by the following steps:
[0007] Step 1: Obtain the time length. The i-th neuromorphic signal stream and the corresponding reference video frame and to Perform voxelization to obtain the i-th neuromorphic voxel. ,in, and These represent the height and width of the neuromorphic voxel, respectively; C represents the number of channels after voxelization. Represent the space of real numbers;
[0008] Step 2: Construct a video reconstruction network based on state-space equations, including: a shallow neuromorphic feature extraction module, a spatial locality enhancement downsampling module, a spatiotemporal locality enhancement module, a spatial locality enhancement upsampling module, and a video output module; and perform... Processing is performed to obtain the i-th reconstructed video frame. ;
[0009] Step 2.1: The shallow neuromorphic feature extraction module utilizes spatial convolution residuals to extract neuromorphic voxels. Feature extraction is performed to obtain the i-th shallow neural mimicry feature. ,in, The number of channels representing superficial neural mimicry features;
[0010] Step 2.2: The spatial locality enhancement downsampling module consists of M cascaded random window downsampling modules, and it is used for shallow neuromorphic features. After processing, the i-th spatially enhanced neural mimicry feature set is output. ;in, This represents the neuromorphic features after local spatial enhancement at level m;
[0011] Step 2.3: The spatiotemporal locality enhancement module consists of N cascaded Mamba modules, and enhances the neuromorphic features of the Mth level spatial locality. After processing, the i-th spatiotemporally localized enhanced neuromorphic feature set is output. ;in, This represents the neural mimicry features after the nth level of spatiotemporal local enhancement;
[0012] Step 2.4: The spatial locality enhancement upsampling module consists of M cascaded random window upsampling modules, and it applies the neuromorphic features after the nth level of spatiotemporal local enhancement. and After processing, the i-th neuromorphic fusion feature set is output. ;in, This represents the m-th level of neural mimicry fusion feature;
[0013] Step 2.5: The video output module utilizes spatial convolution residuals to fuse the M-th level neuromorphic features. After processing, the i-th reconstructed video frame is obtained. ;
[0014] Step 3: Construct a pre-trained fixed-sensor neural network and... and The i-th perceptual similarity index is obtained through processing. ;
[0015] Step 4: Based on and as well as Construct the total loss function of the video reconstruction network based on state-space equations. The Adam optimizer is used to train a video reconstruction network based on state-space equations to update network parameters until the total loss function is reached. The process continues until convergence, thus obtaining the optimal video reconstruction model, which is used to reconstruct the video from the input neuromorphic signal.
[0016] The video reconstruction method driven by neuromorphic signals based on state-space equations described in this invention is also characterized in that each cascaded random window downsampling module in step 2.2 includes: a random window displacement Mamba module, a downsampling module, and a convolutional long short memory network;
[0017] Step 2.2.1: When m=1, the random window displacement Mamba module in the m-th cascaded random window downsampling module is adjusted according to equation (1). The processing yields the encoded features after spatial local enhancement at level m. :
[0018] (1)
[0019] In equation (1), For random window shifting layers, it means that... Randomly divided into a fixed number of patch blocks; It is a multi-layer sensor; For visual Mamba modules, reshape is the reshaping operation; Indicates random displacement. This represents the amount of random displacement in the height direction. This represents the amount of random displacement in the width direction; Indicates uniform distribution. Represents all possible displacements in real space. This represents the spatial dimensions of the random displacement window;
[0020] Step 2.2.2: When m=1, the downsampling module in the m-th cascaded random window downsampling module... The processing yields the spatially enhanced encoded features after the m-th level downsampling. ;
[0021] Step 2.2.3: When m=1, the convolutional long short memory network in the m-th cascaded random window downsampling module... Processing is performed to obtain the neural mimicry features after spatial local enhancement at level m. ;
[0022] Step 2.2.4: When m=2,3,…,M, enhance the neural mimicry features of the (m-1)th level spatial local enhancement. The input is processed in the m-th cascaded random window downsampling module to obtain the M-th level spatially enhanced neuromorphic features. .
[0023] Furthermore, each cascaded Mamba module in step 2.3 includes: a random window displacement Mamba module and a Hilbert curve filled Mamba module;
[0024] Step 2.3.1: When n=1, the nth cascaded random window shifts the Mamba module. After processing, the bottleneck features after spatial local enhancement at the nth level are obtained. ;
[0025] Step 2.3.2: When n=1, the nth cascaded Hilbert curve filling Mamba module is based on equation (2). After processing, the bottleneck features after spatiotemporal local enhancement at the nth level are obtained. ;
[0026] (2)
[0027] In equation (2), This indicates the Hilbert curve fill operation, meaning that... Divide the area into a fixed number of patch blocks according to the filling method of the Hilbert curve; This indicates the position code added when processing the patch block; This indicates an operation that reverses the Hilbert fill method to reshape the patch block;
[0028] Step 2.3.3: When n=2,3,……,N, apply the bottleneck features after spatiotemporal enhancement at level n-1. The input is processed in the nth cascaded Mamba module, resulting in the Nth cascaded Mamba module outputting the bottleneck features after Nth-level spatial local augmentation. .
[0029] Furthermore, each cascaded random window upsampling module in step 2.4 includes: a random window displacement Mamba module and an upsampling module;
[0030] Step 2.4.1: When m=1, and After element-wise addition, the neuromorphic features after local enhancement at level m are obtained. The random window displacement Mamba module in the m-th cascaded random window upsampling module is then processed to obtain the m-th level spatiotemporal neuromorphic fusion feature. ;
[0031] Step 2.4.2: When m=1, the upsampling module in the m-th cascaded random window upsampling module... Processing was performed to obtain the m-th level neuromorphic fusion feature. ;
[0032] Step 2.4.3: When m=2,3,…,M, fuse the features of the (m-1)th level neuromorphic features. and spatially enhanced neuromorphic features at level M-m+1 After element-wise addition, the neuromorphic features after local enhancement at level m are obtained. The data is then input into the m-th cascaded random window upsampling module for processing, resulting in the output of the M-th level neuromorphic fusion feature from the M-th cascaded random window upsampling module. .
[0033] Furthermore, in step 3, the i-th perceptual similarity index is obtained using equation (3). :
[0034] (3)
[0035] In equation (3), k represents the number of layers in the fixed-sensor neural network, and K represents the total number of layers in the fixed-sensor neural network. This represents the weights corresponding to the k-th layer of the fixed-sensory neural network. This represents the k-th layer fixed-sensory neural network pair. The extracted feature vector, This represents the state of the k-th layer of the fixed-sensory neural network. The extracted feature vector, This represents the distance of the i-th feature in the k-th layer of the fixed-sensory neural network.
[0036] Furthermore, in step 4, the i-th loss function is established using equation (4). Thus, the total loss function is obtained. Where I represents the total number of video frames;
[0037] (4)
[0038] In equation (4), For time consistency loss, This indicates the video frame number at which the calculation of temporal consistency loss begins. Indicates hyperparameters, Let be the i-th perceptual loss function, and we have:
[0039] (5).
[0040] The present invention provides an electronic device, including a memory and a processor, wherein the memory is used to store a program that supports the processor in executing the video reconstruction method, and the processor is configured to execute the program stored in the memory.
[0041] The present invention discloses a computer-readable storage medium on which a computer program is stored, wherein the computer program is executed by a processor to perform the steps of the video reconstruction method.
[0042] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0043] 1. This invention addresses the problem of translation invariance loss in existing visual state space models under fixed partitioned windows by designing a random window offset strategy. This strategy improves the performance of neuromorphic data in model video reconstruction tasks by dynamically and randomly adjusting the window position to better model local information in the spatial domain.
[0044] 2. To address the shortcomings of existing visual state-space models in terms of spatiotemporal locality, this invention proposes a Hilbert space-filling curve mechanism specifically designed for neuromorphic camera video reconstruction tasks. This mechanism effectively captures the spatiotemporal relationships in neuromorphic signal data, significantly enhancing the model's ability to model spatiotemporal information. Attached Figure Description
[0045] Figure 1 This is a flowchart of the video reconstruction method based on state-space equations driven by neuromorphism in this invention;
[0046] Figure 2a This is a schematic diagram of the random window displacement Mamba in this invention;
[0047] Figure 2b This is a schematic diagram of Hilbert filling Mamba in this invention;
[0048] Figure 2c This is a schematic diagram of the visual Mamba module in this invention;
[0049] Figure 2d This is a schematic diagram of the bidirectional visual Mamba module in this invention;
[0050] Figure 3 This is a schematic diagram of random window displacement and Hilbert fill in this invention;
[0051] Figure 4a This is a schematic diagram of the U-Net neural network, representing the overall network structure of this invention.
[0052] Figure 4b This is the residual block network structure used in the shallow neural mimicry feature extraction module and the video output module of this invention;
[0053] Figure 5 This is a graph showing the performance results of this invention compared to other different methods on a low-resolution dataset;
[0054] Figure 6 This is a graph showing the performance results of the present invention compared with other different methods on a high-resolution dataset;
[0055] Figure 7 This is a graph showing the performance results of the present invention compared with other different methods on multiple datasets;
[0056] Figure 8 A comparison chart of the computational resource consumption results of the video reconstruction method of the present invention on multiple datasets. Detailed Implementation
[0057] In this embodiment, a video reconstruction method driven by neuromorphic signals based on state-space equations is proposed. Based on the existing visual Mamba model, and combined with the technical characteristics of video reconstruction tasks driven by neuromorphic signals, the visual Mamba model is structurally optimized and its parameters adjusted. This results in an improved model that combines operational efficiency and reconstruction performance in video reconstruction tasks, significantly improving both the temporal continuity and visual effects of the reconstructed video. Specifically, the method includes the following steps: Figure 1 A flowchart of the video reconstruction algorithm based on state-space equations and driven by neuromorphic signals, as presented in this invention, is given.
[0058] Step 1: Obtain the time length. The i-th neuromorphic signal stream and the corresponding reference video frame and to Perform voxelization to obtain the i-th neuromorphic voxel. ,in, and These represent the height and width of the neuromorphic voxel, respectively; C represents the number of channels after voxelization. Represent the real number space; simultaneously, for the input neuromorphic voxels and reference video frames Perform the same preprocessing steps, including cutting, flipping, and rotating.
[0059] Step 2: Construct a video reconstruction network based on state-space equations. The overall structure is as follows: Figure 1 As shown, the overall network architecture is as follows: Figure 4a The U-Net structure shown includes: a shallow neuromorphic feature extraction module, a spatial locality enhancement downsampling module, a spatiotemporal locality enhancement module, a spatial locality enhancement upsampling module, and a video output module; and it further includes... Processing is performed to obtain the i-th reconstructed video frame. ;
[0060] Step 2.1: The superficial neural mimicry feature extraction module utilizes, for example... Figure 4b shown Spatial convolution residuals for neuromorphic voxels Feature extraction is performed to obtain the i-th shallow neural mimicry feature. ,in, The number of channels representing superficial neural mimicry features;
[0061] Step 2.2: The spatial locality enhancement downsampling module consists of M cascaded random window downsampling modules, and it is used for shallow neuromorphic features. After processing, the i-th spatially enhanced neural mimicry feature set is output. ;in, This represents the neuromorphic features after local spatial enhancement at level m.
[0062] Step 2.2.1: When m=1, the random window displacement Mamba module in the m-th cascaded random window downsampling module is adjusted according to equation (1). The processing yields the encoded features after spatial local enhancement at level m. :
[0063] (1)
[0064] In equation (1), For random window shifting layers, it means that... Randomly divided into a fixed number of patch blocks; It is a multi-layer sensor; For visual Mamba modules, reshape is the reshaping operation; Indicates random displacement. This represents the amount of random displacement in the height direction. This represents the amount of random displacement in the width direction; Indicates uniform distribution. Represents all possible displacements in real space. This represents the spatial dimensions of the random displacement window; Figure 2a This is a schematic diagram of random window shifting in Mamba. Figure 2c This is a schematic diagram of the Visual Mamba module VMB. Figure 2d This is a schematic diagram of a bidirectional Mamba module. Figure 3 Part a in the diagram is a schematic diagram of the random window shifting strategy.
[0065] Step 2.2.2: When m=1, the downsampling module in the m-th cascaded random window downsampling module... The processing yields the spatially enhanced encoded features after the m-th level downsampling. ;
[0066] Step 2.2.3: When m=1, the convolutional long short memory network in the m-th cascaded random window downsampling module... Processing is performed to obtain the neural mimicry features after spatial local enhancement at level m. ;
[0067] Step 2.2.4: When m=2,3,…,M, enhance the neural mimicry features of the (m-1)th level spatial local enhancement. The input is processed in the m-th cascaded random window downsampling module to obtain the M-th level spatially enhanced neuromorphic features. .
[0068] Step 2.3: The spatiotemporal locality enhancement module consists of N cascaded Mamba modules, and enhances the neuromorphic features of the Mth level spatial locality. After processing, the i-th spatiotemporally localized enhanced neuromorphic feature set is output. ;in, This represents the neural mimicry features after the nth level of spatiotemporal local enhancement;
[0069] Step 2.3.1: When n=1, the nth cascaded random window shifts the Mamba module. After processing, the bottleneck features after spatial local enhancement at the nth level are obtained. ;
[0070] Step 2.3.2: When n=1, the nth cascaded Hilbert curve filling Mamba module is based on equation (2). After processing, the bottleneck features after spatiotemporal local enhancement at the nth level are obtained. ;
[0071] (2)
[0072] In equation (2), This indicates the Hilbert curve fill operation, meaning that... Divide the area into a fixed number of patch blocks according to the filling method of the Hilbert curve; This indicates the position code added when processing the patch block; This indicates an operation that reverses the Hilbert fill method to reshape the patch block; Figure 2b A schematic diagram of filling the Hilbert curve with Mamba. Figure 3 Part b in the diagram is a schematic diagram of the Hilbert scan curve.
[0073] Step 2.3.3: When n=2,3,……,N, apply the bottleneck features after spatiotemporal enhancement at level n-1. The input is processed in the nth cascaded Mamba module, resulting in the Nth cascaded Mamba module outputting the bottleneck features after Nth-level spatial local augmentation. .
[0074] Step 2.4: The spatial locality enhancement upsampling module consists of M cascaded random window upsampling modules, and performs upsampling on the neuromorphic features after the nth level of spatiotemporal local enhancement. and After processing, the i-th neuromorphic fusion feature set is output. ;in, This represents the m-th level of neural mimicry fusion feature;
[0075] Step 2.4.1: When m=1, and After element-wise addition, the neuromorphic features after local enhancement at level m are obtained. The random window displacement Mamba module in the m-th cascaded random window upsampling module is then processed to obtain the m-th level spatiotemporal neuromorphic fusion feature. ;
[0076] Step 2.4.2: When m=1, the upsampling module in the m-th cascaded random window upsampling module... Processing was performed to obtain the m-th level neuromorphic fusion feature. ;
[0077] Step 2.4.3: When m=2,3,…,M, fuse the features of the (m-1)th level neuromorphic features. and spatially enhanced neuromorphic features at level M-m+1 After element-wise addition, the neuromorphic features after local enhancement at level m are obtained. The data is then input into the m-th cascaded random window upsampling module for processing, resulting in the output of the M-th level neuromorphic fusion feature from the M-th cascaded random window upsampling module. .
[0078] Step 2.5: The video output module utilizes, for example... Figure 4b shown Spatial convolution residuals for M-th level neuromorphic fusion features After processing, the i-th reconstructed video frame is obtained. ;
[0079] Step 3: Construct a perceptual similarity calculation module, including a pre-trained fixed perceptual neural network, and perform... and The data is processed, and the i-th perceptual similarity index is obtained according to equation (3). ;
[0080] (3)
[0081] In equation (3), k represents the number of layers in the fixed-sensor neural network, and K represents the total number of layers in the fixed-sensor neural network. This represents the weights corresponding to the k-th layer of the fixed-sensory neural network. This represents the k-th layer fixed-sensory neural network pair. The extracted feature vector, The loss function represents the loss in the k-th layer of the fixed-sensory neural network. The extracted feature vector, This represents the distance of the i-th feature in the k-th layer of the fixed-sensory neural network.
[0082] Step 4: Construct the total loss function for the video reconstruction network based on state-space equations The Adam optimizer is used to train a video reconstruction network based on state-space equations to update network parameters until the total loss function is reached. The optimal video reconstruction model is obtained by converging until the signal converges. This model is then used to reconstruct the video from the input neuromorphic signal. The i-th loss function is established using equation (4). Thus, the total loss function is obtained. Where I represents the total number of video frames;
[0083] (4)
[0084] In equation (4), For time consistency loss, This indicates the video frame number at which the calculation of temporal consistency loss begins. Indicates hyperparameters, Let be the i-th perceptual loss function, and we have:
[0085] (5)
[0086] In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.
[0087] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.
[0088] To quantitatively evaluate the effectiveness of this invention and verify its validity, comparative experiments were conducted with various deep learning-based neuromorphic signal-driven video reconstruction algorithms, including Convolutional Event Reconstruction Network (E2VID), Fast Event Reconstruction Network (FireNet), Fast Event Reconstruction Network+ (FireNet+), Convolutional Event Reconstruction Network+ (E2VID+), Spatial Adaptive Event Reconstruction Network (SPADE-E2VID), Sparse Event Reconstruction Network (Sparse-E2VID), and Transformer Event Reconstruction Network (ET-Net). Experiments were performed on real-world footage from three publicly available datasets: High Quality Events Dataset (HQF), Robotics Dataset (IJRR), and Vehicle Recording Dataset (MVSEC), aiming to comprehensively evaluate the performance of each method. To objectively measure the effectiveness of the method, the following three performance metrics were selected as evaluation criteria: Mean Squared Error (MSE): measures the pixel-level error between the reconstructed image and the real image; a smaller MSE value indicates higher reconstruction quality. Structural Similarity Index (SSIM): evaluates the similarity of two images in structural features; a SSIM closer to 1 indicates better structural consistency. Learned Perceptual Image Patch Similarity (LPIPS): quantifies visual perceptual differences; a lower LPIPS value indicates better perceptual quality.
[0089] Numerical indicator evaluation consists of three parts:
[0090] The first part is based on error-sensitive image quality assessment, with mean squared error (MSE) as the evaluation criterion.
[0091] (5)
[0092] In equation (5), This represents the total number of pixels in the image. This is represented as the value of the j-th pixel in the output video frame. This represents the value of the j-th pixel in the reference video frame.
[0093] The second part is image quality assessment based on structural similarity, using the Structural Similarity Index (SSIM) as the evaluation criterion.
[0094] (6)
[0095] In equation (6), for The mean, for The mean, for variance for variance for and covariance, and It is a constant.
[0096] The third part is the evaluation of learning-aware image similarity, with the evaluation index being LPIPS, as shown in Equation (3).
[0097] Figure 5 This paper presents quantitative comparison results between the video reconstruction network of this invention and existing methods. On the MSE and SSIM metrics, the proposed video reconstruction network outperforms all existing reconstruction networks based on neuromorphic signals. In terms of the LPIPS metric, the proposed method also outperforms most existing methods. Except on high-quality event datasets, the LPIPS metric is slightly lower than the convolutional event reconstruction network+ method. Overall, the proposed method demonstrates superior performance on MSE, SSIM, and LPIPS metrics, achieving state-of-the-art results on multiple datasets.
[0098] Figures 6-7 The visual contrast effect of the invention on the test set is demonstrated. Figure 6This paper presents a visual comparison of the reconstruction results of the proposed method with all baseline methods on video clips from low-resolution datasets, high-quality event datasets, robot datasets, and dashcam datasets. Furthermore, ground truth (GT) frames are provided for reference. It can be observed that the reconstructed frames from the Fast Event Reconstruction Network+ and Convolutional Event Reconstruction Network+ lack accuracy in brightness, resulting in poor overall visual quality. In contrast, the reconstruction results from the Transformer Event Reconstruction Network are visually more realistic than those from the Fast Event Reconstruction Network+ and Convolutional Event Reconstruction Network+, presenting an effect closer to the real scene. The proposed method further enriches the final reconstruction results, revealing more fine details while significantly reducing the artifacts common in Convolutional Event Reconstruction Network+. Moreover, the image contrast of the reconstructed frames from this invention is very close to that of the ground truth (GT). To further support the comparative analysis of this invention, a series of high-resolution video sequences were captured using a Prophesee EVK4 camera in this embodiment. Figure 7 A series of comparative images are shown, demonstrating that the method proposed in this invention provides superior visual effects with significantly fewer artifacts compared to other methods.
[0099] To verify the efficiency of the network of this invention, this embodiment analyzes its computational complexity compared to other methods, focusing on the total number of parameters and inference time on 3090 GPUs. To make the comparison more meaningful, this invention selects two commonly used resolution types of neuromorphic camera sensors: low resolution... (L) and high resolution (H). Experimental results are as follows: Figure 8 As shown in the table. The number of model parameters is in millions (M), and inference time is in milliseconds (ms). From... Figure 8 It can be seen that the method of this invention achieves a good balance between computational complexity and performance. On the one hand, compared with the Transformer event reconstruction network, the method of this invention not only reduces the number of parameters by half, but also improves performance in processing... The method significantly reduces inference time when dealing with large input sizes. Furthermore, compared to CNN-based methods, the method of this invention achieves a significant performance improvement with only a slight increase in computational cost, fully demonstrating the superiority of the method.
Claims
1. A video reconstruction method driven by neuromorphic signals based on state-space equations, characterized in that, Includes the following steps: Step 1: Obtain the time length. The i-th neuromorphic signal stream and the corresponding reference video frame and to Perform voxelization to obtain the i-th neuromorphic voxel. ,in, and These represent the height and width of the neuromorphic voxel, respectively; C represents the number of channels after voxelization. Represents the space of real numbers; Step 2: Construct a video reconstruction network based on state-space equations, including: a shallow neuromorphic feature extraction module, a spatial locality enhancement downsampling module, a spatiotemporal locality enhancement module, a spatial locality enhancement upsampling module, and a video output module; and perform... Processing is performed to obtain the i-th reconstructed video frame. ; Step 2.1: The shallow neuromorphic feature extraction module utilizes spatial convolution residuals to extract neuromorphic voxels. Feature extraction is performed to obtain the i-th shallow neural mimicry feature. ,in, The number of channels representing superficial neural mimicry features; Step 2.2: The spatial locality enhancement downsampling module consists of M cascaded random window downsampling modules, and it is used for shallow neuromorphic features. After processing, the i-th spatially enhanced neural mimicry feature set is output. ;in, This represents the neuromorphic features after local spatial enhancement at level m; Step 2.3: The spatiotemporal locality enhancement module consists of N cascaded Mamba modules, and enhances the neuromorphic features of the Mth level spatial locality. After processing, the i-th spatiotemporally localized enhanced neuromorphic feature set is output. ;in, This represents the neural mimicry features after the nth level of spatiotemporal local enhancement; Step 2.4: The spatial locality enhancement upsampling module consists of M cascaded random window upsampling modules, and it applies the neuromorphic features after the nth level of spatiotemporal local enhancement. and After processing, the i-th neuromorphic fusion feature set is output. ;in, This represents the m-th level of neural mimicry fusion feature; Step 2.5: The video output module utilizes spatial convolution residuals to fuse the M-th level neuromorphic features. After processing, the i-th reconstructed video frame is obtained. ; Step 3: Construct a pre-trained fixed-sensor neural network and... and The i-th perceptual similarity index is obtained through processing. ; Step 4: Based on and as well as Construct the total loss function of the video reconstruction network based on state-space equations. The Adam optimizer is used to train a video reconstruction network based on state-space equations to update network parameters until the total loss function is reached. The process continues until convergence, thus obtaining the optimal video reconstruction model, which is used to reconstruct the video from the input neuromorphic signal.
2. The video reconstruction method based on state-space equations and driven by neuromorphic signals according to claim 1, characterized in that, Each cascaded random window downsampling module in step 2.2 includes: a random window shifting Mamba module, a downsampling module, and a convolutional long short memory network; Step 2.2.1: When m=1, the random window displacement Mamba module in the m-th cascaded random window downsampling module is adjusted according to equation (1). The processing yields the encoded features after spatial local enhancement at level m. : (1) In equation (1), For random window shifting layers, it means that... Randomly divided into a fixed number of patch blocks; It is a multi-layer sensor; For visual Mamba modules, reshape is the reshaping operation; Indicates random displacement. This represents the amount of random displacement in the height direction. This represents the amount of random displacement in the width direction; Indicates uniform distribution. Represents all possible displacements in real space. This represents the spatial dimensions of a random displacement window; Step 2.2.2: When m=1, the downsampling module in the m-th cascaded random window downsampling module... The processing yields the spatially enhanced encoded features after the m-th level downsampling. ; Step 2.2.3: When m=1, the convolutional long short memory network in the m-th cascaded random window downsampling module... Processing is performed to obtain the neural mimicry features after spatial local enhancement at level m. ; Step 2.2.4: When m=2,3,…,M, enhance the neural mimicry features of the (m-1)th level spatial local enhancement. The input is processed in the m-th cascaded random window downsampling module to obtain the M-th level spatially enhanced neuromorphic features. .
3. The video reconstruction method based on state-space equations and driven by neuromorphic signals according to claim 2, characterized in that, Each cascaded Mamba module in step 2.3 includes: a random window shift Mamba module and a Hilbert curve fill Mamba module; Step 2.3.1: When n=1, the nth cascaded random window shifts the Mamba module. The bottleneck features are obtained after local spatial enhancement at the nth level. ; Step 2.3.2: When n=1, the nth cascaded Hilbert curve filling Mamba module is based on equation (2). After processing, the bottleneck features after spatiotemporal local enhancement at the nth level are obtained. ; (2) In equation (2), This indicates the Hilbert curve fill operation, meaning that... Divide the area into a fixed number of patch blocks according to the filling method of the Hilbert curve; This indicates the position code added when processing the patch block; This indicates an operation that reverses the Hilbert fill method to reshape the patch block; Step 2.3.3: When n=2,3,……,N, apply the bottleneck features after spatiotemporal enhancement at level n-1. The input is processed in the nth cascaded Mamba module, resulting in the Nth cascaded Mamba module outputting the bottleneck features after Nth-level spatial local augmentation. .
4. The video reconstruction method based on state-space equations and driven by neuromorphic signals according to claim 3, characterized in that, Each cascaded random window upsampling module in step 2.4 includes: a random window shifting Mamba module and an upsampling module; Step 2.4.1: When m=1, and After element-wise addition, the neuromorphic features after local enhancement at level m are obtained. The random window displacement Mamba module in the m-th cascaded random window upsampling module is then processed to obtain the m-th level spatiotemporal neuromorphic fusion feature. ; Step 2.4.2: When m=1, the upsampling module in the m-th cascaded random window upsampling module... Processing was performed to obtain the m-th level neuromorphic fusion feature. ; Step 2.4.3: When m=2,3,…,M, fuse the features of the (m-1)th level neuromorphic features. and spatially enhanced neuromorphic features at level M-m+1 After element-wise addition, the neuromorphic features after local enhancement at level m are obtained. The data is then input into the m-th cascaded random window upsampling module for processing, resulting in the output of the M-th level neuromorphic fusion feature from the M-th cascaded random window upsampling module. .
5. The video reconstruction method based on state-space equations and driven by neuromorphic signals according to claim 4, characterized in that, In step 3, the i-th perceptual similarity index is obtained using equation (3). : (3) In equation (3), k represents the number of layers in the fixed-sensor neural network, and K represents the total number of layers in the fixed-sensor neural network. This represents the weights corresponding to the k-th layer of the fixed-sensory neural network. This represents the k-th layer fixed-sensory neural network pair. The extracted feature vector, This represents the state of the k-th layer of the fixed-sensory neural network. The extracted feature vector, This represents the distance of the i-th feature in the k-th layer of the fixed-sensory neural network.
6. The video reconstruction method based on state-space equations and driven by neuromorphic signals according to claim 5, characterized in that, In step 4, the i-th loss function is established using equation (4). Thus, the total loss function is obtained. Where I represents the total number of video frames; (4) In equation (4), For time consistency loss, This indicates the video frame number at which the calculation of temporal consistency loss begins. Indicates hyperparameters, Let be the i-th perceptual loss function, and we have: (5)。 7. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor in executing the video reconstruction method of any one of claims 1-6, and the processor is configured to execute the program stored in the memory.
8. A computer-readable storage medium storing a computer program thereon, characterized in that, The computer program is executed by the processor to perform the steps of any of the video reconstruction methods described in claims 1-6.
Citation Information
Patent Citations
Butterfly image recognition method based on convolutional neural network
CN115631417A
Image processing method, apparatus, and non-transitory computer-readable medium
US20230325974A1