A time step mixed cross-modal pulse neural network training method and system
By constructing a hybrid sequence of static and event modalities and training a spiking neural network using multiple loss functions, the high training difficulty and transfer problems of event cameras were solved, achieving stable and efficient cross-modal transfer and generalization capabilities.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- NINGBO UNIV
- Filing Date
- 2026-03-24
- Publication Date
- 2026-06-23
AI Technical Summary
Existing technologies struggle to effectively train the spiking neural network of an event camera without altering the backbone network, and direct training from scratch is highly difficult. Furthermore, the transfer from static modes to event modes suffers from distribution shift and negative transfer issues.
A time-step hybrid cross-modal spiking neural network training method is adopted. By constructing a hybrid sequence of static and event modalities, and combining labels, multiple loss functions are calculated, including event flow classification loss, hybrid flow classification loss, domain alignment loss, modality-aware loss and hybrid proportional-aware loss, to update the network parameters.
It achieves smooth transfer in the time dimension, reduces gradient noise during training, improves training stability and convergence speed, reduces negative transfer, enhances the model's generalization performance in event vision tasks, and has interpretability and wide applicability.
Smart Images

Figure CN122263975A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence and neuromorphic computing technology, specifically relating to a time-step hybrid cross-modal spiking neural network training method and system. Background Technology
[0002] Event cameras (Dynamic Vision Sensors, DVS) are known for their high temporal resolution, low latency, and sparse event output; however, their weak texture and extreme sparsity make it difficult to train spiking neural networks (SNNs) directly from scratch. On the other hand, data and models based on static modalities (such as RGB) are rich in shape-texture priors, and directly using them for fine-tuning in the event domain often encounters inter-domain distribution shifts, resulting in negative transfer and insufficient generalization ability to event sequences. Existing cross-domain transfer methods often rely on feature alignment or distillation, but these are unstable during training, sensitive to hyperparameters, or require a dedicated teacher network.
[0003] Therefore, there is an urgent need for a general training mechanism that does not require modification of the backbone network and does not rely on external teachers, so that the model can smoothly and controllably transition from the static domain to the event domain in the time dimension, while obtaining static priors and avoiding overfitting and misfitting to the event domain. Summary of the Invention
[0004] The main objective of this invention is to provide a time-step hybrid cross-modal spiking neural network training method and system to overcome the shortcomings of the prior art.
[0005] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a time-step hybrid cross-modal spiking neural network training method, wherein the cross-modal spiking neural network is used to perform event visual recognition, and comprises: For the same sample, a static modal time series and an event modal time series are constructed. The static modal time series includes multiple static modal frames, and the event modal time series includes multiple event frames. According to the target mixing ratio, some of the static modal frames are replaced with event modal frames to construct a mixed modal time series; The event modality time series is input into a spiking neural network for training, and the event stream classification loss is calculated by combining the labels. The mixed-modal time series is input into the same spiking neural network for training, and the mixed-stream classification loss is calculated by combining the labels; For samples of the same type, the centered kernel alignment similarity of the event modality time series and the mixed modality time series in the spiking neural network is calculated, and the domain alignment loss is generated. The spiking neural network is used to determine whether the current frame belongs to the static mode or the event mode, and the modality perception loss is generated using the true mode of the current frame as the label; The spiking neural network is used to predict the predicted mixing ratio of the two modalities of the sample, and the mixing ratio perception loss is calculated using the actual mixing ratio as the label. The parameters of the spiking neural network are updated by fusing the event stream classification loss, hybrid stream classification loss, domain alignment loss, modality perception loss, and hybrid scaling perception loss.
[0006] Secondly, the present invention also provides a time-step hybrid cross-modal spiking neural network training system, wherein the cross-modal spiking neural network is used to perform event visual recognition, and comprises: The sequence construction module is used to construct a static modal time series and an event modal time series for the same sample. The static modal time series includes multiple static modal frames, and the event modal time series includes multiple event frames. The hybrid sequence module is used to replace some of the static modal frames with event modal frames according to a target mixing ratio, thereby constructing a hybrid modal time series; The event stream loss module is used to input the event modality time series into the spiking neural network for training, and calculate the event stream classification loss by combining the labels. The hybrid flow loss module is used to input the hybrid modal time series into the same spiking neural network for training, and calculate the hybrid flow classification loss by combining the labels. The domain alignment loss module is used to calculate the centered kernel alignment similarity of the event modality time series and the mixed modality time series in the spiking neural network for the same type of samples, and generate the domain alignment loss. The perceptual loss module is used to determine whether the current frame belongs to the static mode or the event mode using the spiking neural network, and to generate a mode perceptual loss with the true mode of the current frame as the label. The proportional loss module is used to predict the predicted mixing ratio of the two modalities of the sample using the spiking neural network, and calculate the mixing ratio perception loss with the true mixing ratio as the label. The fusion update module is used to fuse the event stream classification loss, hybrid stream classification loss, domain alignment loss, modality perception loss, and hybrid proportional perception loss to update the parameters of the spiking neural network.
[0007] Compared with the prior art, the beneficial effects of the present invention include at least the following: Time-dimensional smooth transition and low variance optimization: This invention performs a one-time switch along the time axis within a single sample to construct a mixed sequence of "static and event". Compared with batch-level mixing strategies, under reasonable statistical assumptions, the time-step mixing strategy of this invention can significantly reduce the variance of the gradient during training while keeping the expected gradient unchanged, thereby reducing gradient noise and improving the training stability and convergence speed of spiking neural networks.
[0008] Non-intrusive and easy to integrate: This invention does not require modification of the existing SNN or ANN backbone network structure. It only requires design at the input construction and loss function level and can be directly integrated into mainstream spiking neural networks or neural network training pipelines with time dimension. It has low deployment cost and strong versatility.
[0009] Robust cross-modal transfer and negative transfer suppression: By adopting a smooth transition input from static to event in the time dimension, combined with CKA-based regularization domain alignment constraints and MAG and MRP auxiliary supervision, the distribution mismatch between static and event modalities in the high-dimensional feature space can be significantly reduced, alleviating the negative transfer problem caused by direct fine-tuning, and enabling the model to achieve higher generalization performance on event vision tasks.
[0010] High interpretability and controllability: By explicitly setting the mixing ratio and total time step, this invention can precisely control the expected proportion of static and event data in each sample; at the same time, the MAG and MRP tasks enable the model to have interpretable modality awareness at the time and sequence levels, which is beneficial for flexibly adjusting the mixing strategy according to computing power, energy consumption and task requirements.
[0011] Wide range of applications: This invention is not only applicable to event camera classification tasks, but can also be extended to tasks such as event target detection and event segmentation, as well as other cross-modal learning scenarios that include static modalities and temporal event modalities, and has broad application prospects.
[0012] The above description is merely an overview of the technical solution of the present invention. In order to enable those skilled in the art to better understand the technical means of this application and to implement it in accordance with the contents of the specification, the preferred embodiments of the present invention are described below in conjunction with detailed drawings. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1This is a schematic diagram of the overall architecture of a cross-modal spiking neural network training method provided in a typical embodiment of the present invention. Detailed Implementation
[0015] In view of the shortcomings of the prior art, the inventors of this invention, through long-term research and extensive practice, have proposed the technical solution of this invention. The following will further explain and illustrate this technical solution, its implementation process, and its principles.
[0016] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and therefore the scope of protection of the invention is not limited to the specific embodiments disclosed below.
[0017] The purpose of this invention is to provide a time-step mixup (TSM) method for training cross-modal spiking neural networks. This method employs a truncated and replaced mixing approach on the time axis to construct a mixed temporal input that gradually transitions from static to event-based modes. It combines Regularized Domain Alignment (RDA) with two lightweight auxiliary tasks: frame-by-frame Modality-Aware Guidance (MAG) and sample-level Mixup Ratio Perception (MRP) to achieve efficient transfer and alignment of cross-modal knowledge, thereby improving the accuracy, training stability, and convergence speed of event vision tasks.
[0018] For the purposes described above, see Figure 1 As shown, an embodiment of the present invention first provides a time-step hybrid cross-modal spiking neural network training method, wherein the cross-modal spiking neural network is used to perform event visual recognition, and includes the following steps: For the same sample, a static modal time series and an event modal time series are constructed. The static modal time series includes multiple static modal frames, and the event modal time series includes multiple event frames. According to the target mixing ratio, some of the static modal frames are replaced with event modal frames to construct a mixed modal time series; The event modality time series is input into a spiking neural network for training, and the event stream classification loss is calculated by combining the labels. The mixed-modal time series is input into the same spiking neural network for training, and the mixed-stream classification loss is calculated by combining the labels; For samples of the same type, the centered kernel alignment similarity of the event modality time series and the mixed modality time series in the spiking neural network is calculated, and the domain alignment loss is generated. The spiking neural network is used to determine whether the current frame belongs to the static mode or the event mode, and the modality perception loss is generated using the true mode of the current frame as the label; The spiking neural network is used to predict the predicted mixing ratio of the two modalities of the sample, and the mixing ratio perception loss is calculated using the actual mixing ratio as the label. The parameters of the spiking neural network are updated by fusing the event stream classification loss, hybrid stream classification loss, domain alignment loss, modality perception loss, and hybrid scaling perception loss.
[0019] The temporal step mixing method deterministically constructs a mixed sequence. While regularizing features, this process also introduces two types of uncertainty: local ambiguity at each time step regarding which modality generated the current frame; and global uncertainty regarding the proportion of each modality in the entire sequence. TSM naturally incorporates these two types of information during its construction. This invention transforms this into two lightweight auxiliary tasks to enable the model to learn a temporal representation of the mixed temporal structure. MAG provides frame-by-frame supervision signals, ensuring the network maintains consistent statistical properties across modalities, thus acting as local supervision. MRP, on the other hand, helps the temporal features reflect the global mixing ratio, clarifying the proportion of different modalities in the sequence and the timing of their switching.
[0020] The two lightweight tasks have a synergistic effect on improving network performance, as can be seen in the experimental results in the examples below.
[0021] In a typical embodiment, the Transfer_VGG_SNN constructed in this paper employs a VGG-style spiking convolutional network with shared parameters as the backbone feature extractor for unified representation learning of RGB and DVS modalities. This backbone network consists of 8 convolutional modules and 4 average pooling layers. All convolutional layers use 3×3 convolutional kernels, and a 2×2 average pooling layer follows every two convolutional modules. The number of network channels is set to 64, 128, 256, 512, and 512 respectively, resulting in a deep feature representation of size 512×3×3. Since each convolutional module incorporates LIF spiking neurons, the proposed backbone network can simultaneously model the spatial structure and temporal dynamics of the input data, providing effective feature support for subsequent cross-modal transfer and knowledge distillation.
[0022] Of course, the key point of the technical solution provided by this invention lies in the design from the level of input construction and loss function, rather than being limited to a specific part of the backbone network. In fact, the technical solution provided by this invention has good adaptability to a variety of backbone networks.
[0023] In some implementations, the construction process of the mixed-modal time series specifically includes the following steps: Sampling ; structure , making when season ; when season ; in, Indicates switching time steps. This represents the total number of time steps in the static modal time series. This represents the mixed-modal time series. This represents a frame in the mixed-modal time series. Indicates the frame number. This refers to the static modal frame. This refers to the event modal frame.
[0024] In some implementations, the sampling process for the switching time step specifically includes the following steps: set up ∈(0,1); Let the switching time step be The mixed-modal time series follows a truncated geometric distribution with parameter p. Truncating and normalizing this distribution transforms it into a valid probability distribution, enabling the calculation, for a given parameter p, of the expected number of suffix steps to be replaced under this truncated geometric distribution.
[0025] Solve for the parameter p using numerical methods, such that
[0026] in, Indicates the target mixing ratio. Indicates the number of steps in the suffix. This indicates the overall expected replacement.
[0027] In some implementations, the loss function for the event stream classification loss and / or the hybrid stream classification loss is the TET loss function.
[0028] In some implementations, the domain alignment loss is expressed as follows:
[0029] in, This represents the domain alignment loss. The central kernel alignment method (Simon Kornblith, Mohammad Norouzi, Honglak Lee, and Geoffrey Hinton. Similarity of neural network representations revisited. In International conference on machine learning, pages 3519–3529. PM1R, 2019.5) is used to measure the similarity between two networks. , These represent the membrane potentials of the penultimate layer of the spiking neural network at time step t for mixed sample i and event sample j, respectively, where i and j belong to the same category. , Let i and j represent the categories of samples i and j respectively, and Y represent the total set of sample categories.
[0030] In some implementations, the calculation process of the modality sensing loss specifically includes the following steps: During the generation of the mixed-mode time series At the same time, record the true modal label corresponding to each time step t: if Then it is marked as "static mode". It is then marked as "event modality"; Set up a modality discriminant head network, take the intermediate features of each time step as input, and output the predicted probability of the current time step belonging to the static mode or the event mode; Calculate the frame-by-frame cross-entropy loss as the modality-aware loss. .
[0031] In some implementations, the calculation process of the hybrid ratio sensing loss specifically includes the following steps: The actual mixing ratio is calculated as follows: ; The features at each time step are mapped to scalars (quantities with magnitude but no direction) through a regression head, and then averaged over the time dimension to obtain the predicted mixing ratio. ; Calculate the predicted mixing ratio With the actual mixing ratio The mean squared error loss is used to obtain the hybrid proportional sensing loss. .
[0032] In some implementation schemes, the process of obtaining loss through fusion specifically includes the following steps: A floating weight parameter that can be learned over time steps is introduced to weight and combine the domain alignment loss and the event flow classification loss to form a spatiotemporal regularized alignment loss. The floating weight parameter is automatically optimized through backpropagation. The spatiotemporal regularization alignment loss, the hybrid stream classification loss, the modality-aware loss, and the hybrid proportional-aware loss are summed to form the total loss, which is used to update the parameters of the spiking neural network.
[0033] Because event data is dynamic, using only domain alignment loss for spatial feature alignment may result in the loss of important temporal information. Spatiotemporal regularization provides dynamically learnable coefficients for the domain alignment loss; these adaptive coefficients ensure that specific weights are assigned to data features at each time step. To avoid overfitting the model at any given time step, we use the adaptively adjusted event data classification loss at each time step as a regularization term.
[0034] In some implementations, the spatiotemporal regularization alignment loss and total loss are calculated as follows:
[0035]
[0036] in, This represents the spatiotemporal regularization alignment loss. This represents the total loss. This represents the floating weight parameter. σ(·) represents the learnable coefficients of time step t. Specifically, σ(·) represents the sigmoid function that maps it to the range of 0 to 1, and is a commonly used function in the field of computer science. This represents the domain alignment loss. This represents the event stream classification loss. This represents the mixed-stream classification loss. This represents a fixed weight parameter, a type of hyperparameter. Its typical value is between 0 and 1; in a typical embodiment of this invention, it is set to 0.5. However, the range of values for this parameter is not limited to the examples provided in this invention. This represents the modality sensing loss. This represents the perceived loss of the mixing ratio.
[0037] Embodiments of the present invention also provide a time-step hybrid cross-modal spiking neural network training system, wherein the cross-modal spiking neural network is used to perform event visual recognition, and includes: The sequence construction module is used to construct a static modal time series and an event modal time series for the same sample. The static modal time series includes multiple static modal frames, and the event modal time series includes multiple event frames. The hybrid sequence module is used to replace some of the static modal frames with event modal frames according to a target mixing ratio, thereby constructing a hybrid modal time series; The event stream loss module is used to input the event modality time series into the spiking neural network for training, and calculate the event stream classification loss by combining the labels. The hybrid flow loss module is used to input the hybrid modal time series into the same spiking neural network for training, and calculate the hybrid flow classification loss by combining the labels. The domain alignment loss module is used to calculate the centered kernel alignment similarity of the event modality time series and the mixed modality time series in the spiking neural network for the same type of samples, and generate the domain alignment loss. The perceptual loss module is used to determine whether the current frame belongs to the static mode or the event mode using the spiking neural network, and to generate a mode perceptual loss with the true mode of the current frame as the label. The proportional loss module is used to predict the predicted mixing ratio of the two modalities of the sample using the spiking neural network, and calculate the mixing ratio perception loss with the true mixing ratio as the label. The fusion update module is used to fuse the event stream classification loss, hybrid stream classification loss, domain alignment loss, modality perception loss, and hybrid proportional perception loss to update the parameters of the spiking neural network.
[0038] Embodiments of the present invention also provide a readable storage medium storing a computer program that, when run, performs the steps of the training method provided in any of the above embodiments.
[0039] The technical solution of the present invention will be further described in detail below through several embodiments and in conjunction with the accompanying drawings. However, the selected embodiments are only for illustrating the present invention and do not limit the scope of the present invention.
[0040] Example 1 This embodiment provides a time-step hybrid cross-modal spiking neural network training method, including the following steps: 1. Data sequence construction: For each sample, construct a static modal time series. The static modality can be obtained by copying a static image over time (T frames) or by uniformly sampling T frames from a video. Correspondingly, an event modality time series is constructed. The event modality is a framed event graph.
[0041] 2. Set the mixing ratio for each time step: Set target mixing ratio ∈(0,1), given a total number of time steps T, it is expected that there are approximately 10 ... Each suffix time step replaces the static modal frame with an event modal frame to control the blending level and the transition length from static to event.
[0042] 3. Construct the truncated geometric distribution and solve for the switching parameters: Let the switching time step of the mixed sequence be... It follows a truncated geometric distribution with parameter p, and its support set is defined as follows. Furthermore, by truncating and normalizing it to make it a valid probability distribution, we can calculate, for a given p, the expected number of suffix replacement steps under this distribution.
[0043] Solving for the parameter p using numerical methods (such as the bisection method or Newton's method) allows for the following:
[0044] This ensures, statistically speaking, that the mixing ratio is close to the target mixing ratio. .
[0045] 4. Sample and switch time steps according to the truncated geometric distribution and generate a mixed sequence: For each training sample: The sampling switching time step is based on the truncated geometric distribution. ; Constructing mixed sequences ,in when ,make (Using static modal frames); when ,make (Using event modal frames); By using the above method, a time series of the first half of static events and the second half of events is constructed within each sample, thereby achieving modal gradation in the time dimension.
[0046] 5. Forward Propagation and Regularized Domain Alignment (RDA) Pure event modal sequences By inputting the same spiking neural network, the penultimate layer membrane potential representations at each time step are obtained. The event stream classification loss is calculated using the TET loss function, along with the classification output. ; Mixed sequences Simultaneously inputting the data into the network yields the membrane potential representation for the corresponding time step. And use the TET loss function to calculate the mixed-stream classification loss. ; For mixed membrane potential sequences and event membrane potential sequences of the same category of samples, the similarity of centered kernel alignment (CKA) is calculated by time step, and the domain alignment loss is calculated. ; .
[0047] in , They represent the membrane potential of the penultimate layer of mixed sample i at time step t and the membrane potential of the penultimate layer of event sample j at time step t, respectively, and i and j belong to the same category.
[0048] To balance cross-modal alignment and event domain discrimination capabilities, learnable weight parameters are introduced to adjust the "CKA alignment loss" and "event flow classification loss". "Weighing and combining these factors together, we obtain the spatiotemporal regularized alignment loss:" .
[0049] 6. Frame-by-frame modality-aware supervised MAG: In generating mixed sequences At the same time, this invention simultaneously records the true modal label corresponding to each time step: if Then it is marked as "static mode". This is then labeled as "event mode"; using the intermediate features of each time step as input, a mode discriminator network is set up to output the predicted probability that the current time step belongs to the static mode or the event mode. The mode perception loss is defined using frame-by-frame cross-entropy loss. This encourages networks to explicitly distinguish and perceive the timing of mode switching in time series.
[0050] 7. Sample-level mixed-proportion sensing supervised MRP: For a given mixed sample, the number of its true static frames is The corresponding actual mixing ratio is By mapping the features at each time step to a scalar through a regression head and averaging them over the time dimension, the predicted value of the mixing ratio is obtained. ; in true proportion For supervision, a hybrid proportional sensing loss is constructed using mean squared error loss (MSE). This guides the network to perceive the overall proportion of static and event samples globally.
[0051] 8. Total Loss and Parameter Updates: Construct the total loss function by combining the above loss terms:
[0052] in To adjust the weight coefficients of the alignment strength in the regularization domain, stochastic gradient descent or adaptive optimization algorithms (such as Adam, RMSProp, etc.) are used to perform end-to-end backpropagation updates on the parameters of the spiking neural network and each auxiliary head network until training converges.
[0053] As shown in Table 1, on multiple mainstream event vision datasets, the TMKT method proposed in this invention achieves state-of-the-art accuracy performance under the same network architecture and time step settings. Compared with existing methods such as transfer learning and efficient training, this method demonstrates superior performance in different task scenarios.
[0054] Table 1. Comparison of test performance of different methods
[0055] The model sources for the different comparison methods in the table above are shown below: [1] Yuhang Li, Youngeun Kim, Hyoungseob Park, Tamar Geller, andPriyadarshini Panda. Neuromorphic data augmentation for training spikingneural networks. In European Conference on Computer Vision, pages 631–649. Springer, 2022. [2] Guobin Shen, Dongcheng Zhao, and Yi Zeng. Eventmix: An efficient data augmentation strategy for event-based learning. Information Sciences, 644:119170, 2023. [3] Shikuang Deng, Yuhang Li, Shanghang Zhang, and Shi Gu. Temporalefficient training of spiking neural network via gradient re-weighting. arXivpreprint arXiv:2202.11946, 2022. [4] Rui-Jie Zhu, Malu Zhang, Qihang Zhao, Haoyu Deng, Yule Duan, andLiang-Jian Deng. Tcja-snn: Temporal-channel joint attention for spikingneural networks. IEEE Transactions on Neural Networks and Learning Systems,36(3):5112–5125, 2024. [5] Yiting Dong, Dongcheng Zhao, and Yi Zeng. Temporal knowledgesharing enables spiking neural network learning from past and future. IEEETrans. Artif. Intell., 5(7):3524–3534, 2024. [6] Dongcheng Zhao, Guobin Shen, Yiting Dong, Yang Li, and Yi Zeng.Improving stability and performance of spiking neural networks throughenhancing temporal consistency. Pattern Recognit., 159:111094, 2025. [7] Qiugang Zhan, Guisong Liu, Xiurui Xie, Ran Tao, Malu Zhang, andHuajin Tang. Spiking transfer learning from rgb image to neuromorphic eventstream. IEEE Transactions on Image Processing, 2024. [8] Xiang He, Dongcheng Zhao, Yang Li, Guobin Shen, Qingqun Kong, andYi Zeng. An efficient knowledge transfer strategy for spiking neural networks from static to event domain. In Proceedings of the AAAI Conference on Artificial Intelligence, pages 512–520, 2024. [9] Shuhan Ye, Yuanbin Qian, Chong Wang, Sunqi Lin, Jiazhen Xu, Jiangbo Qian, and Yuqi Li. Cross knowledge distillation between artificial and spiking neural networks, 2025.
[10] Henri Rebecq, René Ranftl, Vladlen Koltun, and DavideScaramuzza. High speed and high dynamic range video with an event camera. IEEE transactions on pattern analysis and machine intelligence, 43(6):1964–1980, 2019.
[11] Yang Li, Yiting Dong, Dongcheng Zhao, and Yi Zeng. N-omniglot, alarge-scale neuromorphic dataset for spatio-temporal sparse few-shotlearning. Scientific Data, 9(1):746, 2022. In addition, the embodiments of the present invention also conducted conditional knockout experiments to observe the impact of two lightweight tasks on model performance, and the results are shown in Table 2 below.
[0056] Table 2 Performance comparison of models with different structures
[0057] It is evident that combining these two auxiliary tasks can help the network achieve optimal performance.
[0058] Based on the above embodiments, it is clear that the time-step hybrid cross-modal spiking neural network training method and system proposed in this invention are applicable to various event vision tasks and their deployment on neuromorphic chips. They have advantages such as stable training, simple implementation, and compatibility with mainstream backbone networks. They can be widely used in scenarios that require high dynamic range and low latency perception, such as intelligent security, autonomous driving, robotics, augmented reality (AR) / virtual reality (VR), and industrial inspection, and have good industrial application prospects.
[0059] It should be understood that the above embodiments are merely illustrative of the technical concept and features of the present invention, and are intended to enable those skilled in the art to understand the content of the present invention and implement it accordingly. They should not be construed as limiting the scope of protection of the present invention. All equivalent changes or modifications made in accordance with the spirit and essence of the present invention should be covered within the scope of protection of the present invention.
Claims
1. A time-step hybrid cross-modal spiking neural network training method, wherein the cross-modal spiking neural network is used to perform event visual recognition, characterized in that, include: For the same sample, a static modal time series and an event modal time series are constructed. The static modal time series includes multiple static modal frames, and the event modal time series includes multiple event frames. According to the target mixing ratio, some of the static modal frames are replaced with event modal frames to construct a mixed modal time series; The event modality time series is input into a spiking neural network for training, and the event stream classification loss is calculated by combining the labels. The mixed-modal time series is input into the same spiking neural network for training, and the mixed-stream classification loss is calculated by combining the labels; For samples of the same type, the centered kernel alignment similarity of the event modality time series and the mixed modality time series in the spiking neural network is calculated, and the domain alignment loss is generated. The spiking neural network is used to determine whether the current frame belongs to the static mode or the event mode, and the modality perception loss is generated using the true mode of the current frame as the label; The spiking neural network is used to predict the predicted mixing ratio of the two modalities of the sample, and the mixing ratio perception loss is calculated using the actual mixing ratio as the label. The parameters of the spiking neural network are updated by fusing the event stream classification loss, hybrid stream classification loss, domain alignment loss, modality perception loss, and hybrid scaling perception loss.
2. The cross-modal spiking neural network training method according to claim 1, characterized in that, The construction process of the mixed-modal time series specifically includes: Sampling ; structure , making when season ; when season ; in, Indicates switching time steps. This represents the total number of time steps in the static modal time series. This represents the mixed-modal time series. This represents a frame in the mixed-modal time series. Indicates the frame number. This refers to the static modal frame. This refers to the event modal frame.
3. The cross-modal spiking neural network training method according to claim 2, characterized in that, The sampling process for the switching time step specifically includes: set up ∈(0,1); Let the switching time step be The mixed-modal time series follows a truncated geometric distribution with parameter p. Truncating and normalizing this distribution transforms it into a valid probability distribution, enabling the calculation, for a given parameter p, of the expected number of suffix steps to be replaced under this truncated geometric distribution. ; Solve for the parameter p using numerical methods, such that ; in, Indicates the target mixing ratio. Indicates the number of steps in the suffix. This indicates the overall expected replacement.
4. The cross-modal spiking neural network training method according to claim 1, characterized in that, The loss function for the event stream classification loss and / or hybrid stream classification loss is the TET loss function.
5. The cross-modal spiking neural network training method according to claim 1, characterized in that, The domain alignment loss is expressed as follows: ; in, This represents the domain alignment loss. Indicates the central kernel alignment method, , These represent the membrane potentials of the penultimate layer of the spiking neural network at time step t for mixed sample i and event sample j, respectively, where i and j belong to the same category. , Let i and j represent the categories of samples i and j respectively, and Y represent the total set of sample categories.
6. The cross-modal spiking neural network training method according to claim 1, characterized in that, The calculation process of the modality sensing loss specifically includes: During the generation of the mixed-mode time series At the same time, record the true modal label corresponding to each time step t: if Then it is marked as "static mode". It is then marked as "event modality"; Set up a modality discriminant head network, take the intermediate features of each time step as input, and output the predicted probability of the current time step belonging to the static mode or the event mode; Calculate the frame-by-frame cross-entropy loss as the modality-aware loss. .
7. The cross-modal spiking neural network training method according to claim 1, characterized in that, The calculation process for the hybrid ratio sensing loss specifically includes: The actual mixing ratio is calculated as follows: ; The features at each time step are mapped to scalars through a regression head, and then averaged over the time dimension to obtain the predicted mixing ratio. ; Calculate the predicted mixing ratio With the actual mixing ratio The mean squared error loss is used to obtain the hybrid proportional sensing loss. .
8. The cross-modal spiking neural network training method according to claim 1, characterized in that, Specifically, it includes: A floating weight parameter that can be learned over time steps is introduced to weight and combine the domain alignment loss and the event flow classification loss to form a spatiotemporal regularized alignment loss. The floating weight parameter is automatically optimized through backpropagation. The spatiotemporal regularization alignment loss, the hybrid stream classification loss, the modality-aware loss, and the hybrid proportional-aware loss are summed to form the total loss, which is used to update the parameters of the spiking neural network.
9. The cross-modal spiking neural network training method according to claim 8, characterized in that, The calculation methods for the spatiotemporal regularization alignment loss and the total loss are expressed as follows: ; ; in, This represents the spatiotemporal regularization alignment loss. This represents the total loss. This represents the floating weight parameter. This represents the learnable coefficients at time step t. This represents the domain alignment loss. This represents the event stream classification loss. This represents the mixed-stream classification loss. Indicates a fixed weight parameter. This represents the modality sensing loss. This represents the perceived loss of the mixing ratio.
10. A time-step hybrid cross-modal spiking neural network training system, wherein the cross-modal spiking neural network is used to perform event visual recognition, characterized in that, include: The sequence construction module is used to construct a static modal time series and an event modal time series for the same sample. The static modal time series includes multiple static modal frames, and the event modal time series includes multiple event frames. The hybrid sequence module is used to replace some of the static modal frames with event modal frames according to a target mixing ratio, thereby constructing a hybrid modal time series; The event stream loss module is used to input the event modality time series into the spiking neural network for training, and calculate the event stream classification loss by combining the labels. The hybrid flow loss module is used to input the hybrid modal time series into the same spiking neural network for training, and calculate the hybrid flow classification loss by combining the labels. The domain alignment loss module is used to calculate the centered kernel alignment similarity of the event modality time series and the mixed modality time series in the spiking neural network for the same type of samples, and generate the domain alignment loss. The perceptual loss module is used to determine whether the current frame belongs to the static mode or the event mode using the spiking neural network, and to generate a mode perceptual loss with the true mode of the current frame as the label. The proportional loss module is used to predict the predicted mixing ratio of the two modalities of the sample using the spiking neural network, and calculate the mixing ratio perception loss with the true mixing ratio as the label. The fusion update module is used to fuse the event stream classification loss, hybrid stream classification loss, domain alignment loss, modality perception loss, and hybrid proportional perception loss to update the parameters of the spiking neural network.