Small sample learning recognition method based on spiking neural network
By combining the pulse neural network method of self-feature extraction module and cross feature comparison module, feature representation is optimized and power consumption is reduced, and the calculation cost of small sample learning and feature capture difficulties in deep neural networks in resource-constrained environments is solved, thereby achieving efficient small sample learning.
Patent Information
- Application Number
- CN202510336778.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-07-04
AI Technical Summary
The existing deep neural networks have high computational costs and high energy consumption in small sample learning, which limits their application in resource-constrained environments. In addition, pulsed neural networks perform poorly when capturing complex spatiotemporal features and cross-class comparisons, making it difficult to achieve efficient classification in fields such as rare species recognition and medical image analysis.
Combining the self-feature extraction module and the cross feature comparison module, feature representation is extracted through the pulsed neural network, feature representation is optimized and power consumption is reduced, and the model performance is optimized using time efficiency training and comparison loss.
While ensuring high-performance classification, it significantly reduces power consumption, improves the quality and robustness of features, is suitable for resource-constrained application scenarios, and achieves fast and efficient small sample learning.
Smart Images

Figure CN120259756A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of rare species identification, medical image analysis, etc., and relates to a few-shot learning and recognition method based on a spiking neural network, and particularly relates to a few-shot learning framework that optimizes feature representation and reduces power consumption through a self-feature extraction module and a cross-feature comparison module. Background Art
[0002] Few-shot Learning (FSL) is a technology that can perform effective classification with only a small amount of labeled data. Traditional deep neural networks (DNNs) perform well in few-shot learning, but their high computational cost and large energy consumption limit their application in resource-constrained environments. Spiking neural networks (SNNs), due to their event-driven characteristics and low energy consumption, are particularly suitable for processing sparse and dynamic data, but still have difficulties in capturing complex spatio-temporal features and making cross-class comparisons.
[0003] Existing few-shot learning methods mainly rely on artificial neural networks (ANNs). Although these methods have achieved good classification performance, their high energy consumption and computational cost limit their application in embedded systems and Internet of Things devices. Therefore, how to utilize the low energy consumption characteristics of spiking neural networks to improve the performance and efficiency of few-shot learning has become a current research hotspot.
[0004] The main challenge of few-shot learning lies in how to effectively learn and generalize with limited labeled data. Traditional deep learning methods usually require a large amount of labeled data to train the model, but in practical applications, it is often unrealistic to obtain a large amount of labeled data. Especially in some specific fields, such as medical image analysis, rare species identification, etc., the cost of obtaining labeled data is extremely high. Therefore, how to still maintain a high classification performance with a small amount of labeled data is the core issue of few-shot learning.
[0005] A spiking neural network (SNN) is a computational model that simulates the working mode of the biological nervous system, and has the characteristics of event-driven, low energy consumption, and high parallelism. Compared with traditional deep neural networks, SNNs perform well in processing sparse and dynamic data, especially when processing time series data and real-time data streams, and can significantly reduce the computational cost and energy consumption. In addition, the neurons of SNNs are only activated when they receive sufficient input spikes, and this characteristic gives SNNs a natural advantage in processing few-shot data.
[0006] Although SNNs show potential in few-shot learning, they still face some challenges in practical applications. First, SNNs perform poorly in capturing complex spatio-temporal features, especially when dealing with high-dimensional data, making it difficult to effectively extract and utilize features. Second, when conducting cross-class comparisons, SNNs often require complex training strategies and a large amount of computational resources, which limits their application in resource-constrained environments. In addition, the performance of existing SNN models in few-shot learning still cannot compare with traditional deep learning methods, especially when dealing with complex tasks, the classification accuracy is relatively low.
[0007] To overcome the above challenges, the current research hotspots mainly focus on the following aspects: one is how to design a more efficient SNN architecture to improve its ability to capture complex spatio-temporal features; the second is how to optimize the training strategy of SNNs to reduce computational costs and energy consumption; the third is how to combine the advantages of traditional deep learning methods to improve the performance of SNNs in few-shot learning. In addition, with the rapid development of the Internet of Things and embedded systems, how to achieve efficient few-shot learning in resource-constrained environments has also become an important research direction. Summary of the Invention
[0008] The present invention discloses a few-shot learning method based on spiking neural networks, which optimizes feature representation and reduces power consumption by combining a self-feature extraction module and a cross-feature comparison module.
[0009] The technical solution of the present invention:
[0010] A few-shot learning recognition method based on spiking neural networks, the steps are as follows:
[0011] (1) Basic feature extraction: Use a spiking neural network (SNN) as the backbone network to extract the feature representation of the input image, providing a basic feature map for the subsequent self-feature extraction module;
[0012] Use VGGSNN as the backbone network of the spiking neural network to extract the feature representation F0 of the input image;
[0013] (2) Self-feature extraction module (SFE): Perform self-correlation analysis on the feature representation F0 extracted in step (1), and enhance the internal correlation of the image by calculating the spatial dependence relationship between each pixel point in the feature representation F0 to generate self-correlation features;
[0014] The self-feature extraction module includes the following steps: 1) Given a feature representation F0 ∈ R T×C×H×W , through the unfolding operation, each pixel position x ∈ [1, H] × [1, W] in T×C is expanded in the U and V dimensions to become F’ ∈ R T×C×H×W×U×V; 2) Apply a convolutional block that follows the computationally efficient bottleneck structure to obtain the autocorrelation pattern in F', i.e., a 1×1 convolutional layer for reducing the number of channels, two 3×3 convolutional layers for transformation, and a 1×1 convolutional layer for restoring the number of channels; gradually aggregate the local correlation patterns of F' through the convolutional block to generate autocorrelation features; 3) Insert a LIF neuron layer behind the convolutional block as the activation function to obtain the final feature representation F1; 4) Apply a residual structure based on F0 and F1 to combine the representations of the two patterns to obtain F; based on the following formula:
[0015] F = F0 + F1 (1)
[0016] where F0 is the feature representation after the spiking neural network, and F1 is the feature representation after the self-feature extraction module;
[0017] (3) Cross-Feature Contrast Module (CFC): For steps (1) and (2), extract the features of the support set and the query set through a weight-sharing structure composed of a spiking neural network and a self-feature extraction module, where the support set is the input image containing known classification labels, used to learn and define the feature representations of different classes; the query set refers to the input image to be classified, and its class membership is determined by comparing with the samples in the support set; then cross-compare the support set and the query set, and generate a joint attention map by calculating the similarity and difference between the support set and the query set, to more accurately identify and emphasize the feature regions crucial for classification decisions;
[0018] The cross-feature contrast module includes the following steps: 1) Construct a four-dimensional cross-correlation tensor C ∈
[0019] R H1×W1×H1×W1 , where H1 and W1 are the values of H and W of the input image after being transformed by the spiking neural network and the self-feature extraction module; 2) Adopt a convolutional matching process and use a 4D convolution with a matching kernel to further refine the tensor, specifically composed of two 4D convolutional layers; the first convolutional layer generates multiple correlation tensors with multiple matching kernels, increasing the number of channels to C1, and the second convolutional layer aggregates the generated multiple correlation tensors into a single 4D cross-correlation tensor; 3) Generate the corresponding attention maps A q and A s of the query set and the support set, revealing the relationship between the query set and the support set; based on the following formula:
[0020]
[0021] where, x q and x s are the positions on the feature map respectively, γ is the temperature factor, and C(.) is the cross-correlation tensor;
[0022] (4) Loss calculation: By combining the Time Efficiency Training Loss (TET Loss) and the contrastive loss (InfoNCE Loss), the classification performance of the deep learning network model is optimized by simultaneously optimizing the loss on the time series and the contrastive loss between features of the deep learning network model composed of steps (1), (2), and (3).
[0023] Combining the Time Efficiency Training Loss and the contrastive loss, the total loss is calculated by the following formula:
[0024]
[0025] L total = λL TET +(1 - λ)L info (5)
[0026] where T is the time step; L CE is the cross-entropy loss; F s , F q are the features of the support set and the query set after passing through the self-feature extraction module respectively; A s , A q refer to the attention maps corresponding to the support set and the query set respectively; y is the sample label; sim is the cosine similarity; τ is the scalar temperature factor; λ is a hyperparameter used to balance the two losses;
[0027] (5) Few-shot classification: According to the loss calculation result in step (4), the query set is classified to output the final classification result;
[0028] The few-shot classification step classifies the query set into the closest support set category by calculating the similarity between the query set and the support set.
[0029] Advantages of the present invention: The present invention proposes a few-shot learning method based on spiking neural networks. By combining the self-feature extraction module and the cross-feature contrast module, the feature representation is optimized and the power consumption is significantly reduced. While ensuring high-performance classification, this method utilizes the time dynamic characteristics of spiking neural networks to effectively improve the quality and robustness of features, and is particularly suitable for resource-constrained application scenarios. Finally, fast and efficient few-shot learning is achieved, with broad application potential. Description of the Drawings
[0030] Figure 1 is the overall architecture diagram of the method of the present invention.
[0031] Figure 2 is the visualization result of the spiking activity of the method of the present invention on the CUB dataset.
[0032] Figure 3This is the t-SNE visualization result of the method of the present invention on the N-Omniglot dataset. Detailed implementation manners
[0033] The following further describes the detailed implementation manners of the present invention in combination with the accompanying drawings and technical solutions.
[0034] I. Dataset preprocessing
[0035] The following datasets are used for experiments: N Omniglot, CUB-200-2011, and miniImageNet. At the same time, as a few-shot learning study, it is necessary to ensure that the training set and the test set are different, that is, their intersection is empty. N-Omniglot is a neuromorphic dataset constructed based on the original Omniglot dataset, which consists of 1,623 handwritten characters from 50 different languages. Each character has only 20 different samples. CUB-200-2011 is a dataset focused on fine-grained classification of birds and is widely used in the field of few-shot learning. It has a total of 200 categories, of which 100 are used for training, 50 for validation, and 50 for testing. miniImageNet is a dataset derived from ImageNet, containing a total of 60,000 images, which are evenly distributed among 100 different object categories. Among these categories, 64 are designated for training, 16 for validation, and 20 are reserved for testing.
[0036] II. Feature extraction
[0037] The backbone network used is VGGSNN, specifically the spiking form of VGG16, with a total of 8 Conv-BN-LIF layers. The backbone network we used is relatively simple because an overly complex backbone network is prone to overfitting and is not conducive to the generalization ability of few-shot learning. The average pooling layer is not applied because the average pooling operation will lose spatial information. A key feature of SNNs is that they can operate sparsely, that is, neurons will generate spikes only when the input stimulus exceeds a certain threshold. Avoiding the use of average pooling can help maintain this sparsity, which may be more efficient in hardware implementation and reduce unnecessary computations. For the input format, the static dataset is replicated for T time steps to meet the input requirements of the spiking backbone network. For neuromorphic datasets, due to their natural T dimension, not many operations need to be performed. Temporarily call the preliminary features obtained through the spiking backbone network F0. Use VGGSNN as the backbone network of the spiking neural network to extract the basic feature representation of the input image. VGGSNN consists of 8 convolutional-batch normalization-leaky integrate-and-fire (Conv-BN-LIF) layers and can effectively extract the spatio-temporal features of images.
[0038] III. Self-Feature Extraction Module (SFE)
[0039] The SFE module pays more attention to the information inside the image and provides reliable input for the CFC module. Given a feature representation F0 ∈ R T×C×H×W , where T is the time step, C is the channel, and H×W represents the spatial resolution. First, we use the unfold operation to unfold each position x ∈ [1, H]×[1, W] in T×C into U and V dimensions, becoming F’ ∈ R T×C×H×W×U×V . We use temporal-channel autocorrelation to preserve the rich semantics of the feature representation for classification, thereby suppressing appearance variations and revealing structural patterns. Then, we apply convolutional blocks following a computationally efficient bottleneck structure to obtain the autocorrelation patterns in F’, i.e., a 1×1 convolutional layer to reduce the number of channels and two 3×3 convolutional layers for transformation, and a 1×1 convolutional layer to restore the size of the number of channels. This series of convolutional operations gradually aggregates local correlation patterns without padding, reducing the U×V dimension to 1×1. After that, a LIF neuron layer is inserted behind the convolutional block as the activation function.
[0040] F = F0 + F1 (1)
[0041] For the entire SFE module, the feature dimension remains unchanged as T×C×H×W. This feature extraction is complementary to the feature F obtained after the spiking backbone. Therefore, we combine the representations of the two modes and apply a residual structure such that the feature sent to the CFC module is the sum of the two, which strengthens the representation of the relational features for the basic features and helps few-shot learning better understand "observing oneself" in the image. This method helps identify intra-class features and helps generalize to unseen target categories.
[0042] IV. Cross-Feature Contrast Module (CFC)
[0043] For steps (1) and (2), the features of the support set and the query set are extracted through a weight-sharing structure composed of a spiking neural network and the self-feature extraction module. The support set is the input image containing known classification labels, which is used to learn and define the feature representations of different classes; the query set refers to the input image to be classified, and its class membership is determined by comparing with the samples in the support set; then, cross-contrast is performed on the support set and the query set, and joint attention maps are generated by calculating the similarities and differences between the support set and the query set, more precisely identifying and emphasizing the feature regions crucial for classification decisions;
[0044] The CFC module takes the input pairs F s and F q of the support set and the query set, and generates corresponding attention maps A q and A s。We first take the average over T dimensions for convenience in subsequent operations. Then, we use a pointwise convolutional layer to transform the query and support representations F q and F s into a more compact representation, reducing the channel dimension C to C′. We construct a four-dimensional cross-correlation tensor C ∈ R H1×W1×H1×W1 , where H1 and W1 refer to the values of H and W of the input image after being transformed by the spiking neural network and the self-feature extraction module. Since the dimensions H and W have been reduced after the convolutional layer of the spiking neural network, the memory usage of the cross-correlation tensor is not particularly large. However, due to some large appearance variations in the few-shot learning setting, we adopt a convolutional matching process. We further refine the tensor using 4D convolution with matching kernels, which specifically consists of two 4D convolutional layers. The first convolution generates multiple correlation tensors with multiple matching kernels, increasing the number of channels to C1, and the second convolution aggregates them into a single 4D cross-correlation tensor. From the refined cross-correlation tensor C, we generate the joint attention maps A q and A s , which reveal the relevant content between the query set and the support set.
[0045]
[0046] where x is the position on the feature map, γ is the temperature factor, and C(.) is the cross-correlation tensor. Note that the value A q (x q ) can be interpreted as the ratio of the matching score of x q to the average probability that the position on the query image matches the position on the support image. Similarly, the attention map for the support is calculated by switching the query and the support in Equation 2. These joint attention maps improve the accuracy of few-shot classification by cross-correlating patterns and adjusting the "important positions of joint attention" according to the images given at test time.
[0047] V. Training Strategy
[0048] We train the network in a single-stage manner, combining two losses to guide the model for accurate classification: the TET-based loss and the contrast-based loss. First, we append a fully-connected classification layer after F to calculate L TET , which guides the model to correctly classify the query of class c ∈ C train . The contrast-based metric loss L infoCalculate the cosine similarity between the query and the support prototype embedding, and finally calculate the infoNCE loss to map the query embedding to the support embedding of the same class. During inference, the query class is predicted as the class of the closest support set. Since SNN has an additional time dimension compared to ANN, we use the TET loss to train our spiking neural network. It has been proven that the TET loss is effective for spiking neural networks. The calculation of the TET loss is as follows:
[0049]
[0050] where T is the time step, and L CE represents the cross-entropy loss, y is the sample label, and F q is the feature of the query set after passing through the self-feature extraction module.
[0051] Next, we use pooling to obtain the final feature representation based on the contrastive metric loss. First, divide the support set into positive and negative classes according to the labels so that the network can better learn the correct classes, and then calculate the following contrastive loss:
[0052]
[0053] where sim(·,·) is the cosine similarity, τ is the scalar temperature factor, Fs and F q are the features of the support set and the query set after passing through the self-feature extraction module respectively; A s , A q refer to the corresponding attention maps of the support set and the query set respectively. During the inference process, the class of the query is predicted as the class of the closest prototype. The final loss function combines these two losses, where λ is the hyperparameter that balances the loss terms:
[0054] L total = λL TET + (1 - λ)L info (5)
[0055] Example 1
[0056] On the N-Omniglot dataset, the method of the present invention achieved an accuracy of 95.3% in the 5-way 1-shot setting, significantly outperforming the existing MAML and Siamese methods. On the CUB and miniImageNet datasets, the method of the present invention also achieved comparable performance to the ANNs methods.
[0057] Table 1 Performance comparison of different SNN-based models on the N-Omniglot dataset:
[0058]
[0059] Table 2 Performance results of the proposed model on the static few-shot learning dataset:
[0060]
[0061] Table 3 Ablation experiments on the SFE and CFC modules:
[0062]
[0063]
Claims
1. A few-shot learning recognition method based on spiking neural networks, characterized in that, The steps are as follows: (1) Basic feature extraction: Use a spiking neural network as the backbone network to extract the feature representation of the input image, providing a basic feature map for the subsequent self-feature extraction module; Use VGGSNN as the backbone network of the spiking neural network to extract the feature representation F0 of the input image; (2) Self-feature extraction module: Perform self-correlation analysis on the feature representation F0 extracted in step (1). By calculating the spatial dependence relationship between each pixel point in the feature representation F0, enhance the correlation within the image and generate self-correlation features; The self-feature extraction module includes the following steps: 1) Given a feature representation F0 ∈ R T×C×H×W , through an unfolding operation, each pixel position x ∈ [1, H] × [1, W] in T × C is expanded in the U and V dimensions to become F’ ∈ R T×C×H×W×U×V ; 2) Apply a convolutional block following a computationally efficient bottleneck structure to obtain the self-correlation pattern in F’, i.e., a 1×1 convolutional layer for reducing the number of channels, two 3×3 convolutional layers for transformation, and a 1×1 convolutional layer for restoring the number of channels; gradually aggregate the local correlation patterns of F’ through the convolutional block to generate self-correlation features; 3) Insert a LIF neuron layer behind the convolutional block as an activation function to obtain the final feature representation F1; 4) Apply a residual structure based on F0 and F1 to combine the representations of the two patterns to obtain F; Based on the following formula: F = F0 + F1 (1) where F0 is the feature representation after passing through the spiking neural network, and F1 is the feature representation after passing through the self-feature extraction module; (3) Cross-feature contrast module: For steps (1) and (2), use a weight-sharing structure composed of a spiking neural network and a self-feature extraction module to extract the features of the support set and the query set. The support set is the input image containing known classification labels, used to learn and define the feature representations of different classes; the query set refers to the input image to be classified, and its class membership is determined by comparing with the samples in the support set; then perform cross-contrast on the support set and the query set, and generate a joint attention map by calculating the similarity and difference between the support set and the query set, more accurately identifying and emphasizing the feature regions crucial for classification decisions; The cross - feature comparison module includes the following steps: 1) Construct a four - dimensional cross - correlation tensor C ∈ R H1×W1×H1×W1 , where H1 and W1 are the values of H and W of the input image after being transformed by the spiking neural network and the self - feature extraction module; 2) Adopt a convolution matching process and use a 4D convolution with a matching kernel to further refine the tensor, which is specifically composed of two 4D convolutional layers; The first convolutional layer generates multiple cross - correlation tensors with multiple matching kernels, increasing the number of channels to C1, and the second convolutional layer aggregates the generated multiple cross - correlation tensors into a single 4D cross - correlation tensor; 3) Generate the attention maps A q and A s , revealing the relationship between the query set and the support set; Based on the following formula: where x q and x s respectively refer to positions on the feature map, γ is the temperature factor, and C(.) is the cross-correlation tensor; (4) Loss calculation: Combine the time-efficiency training loss and the contrast loss, and optimize the classification performance of the deep learning network model by simultaneously optimizing the loss based on the time series and the contrast loss between features of the deep learning network model composed of steps (1), (2), and (3); Combine the time-efficiency training loss and the contrast loss, and calculate the total loss through the following formula: L total = λL TET +(1 - λ)L info (5) where T is the time step; L CE is the cross-entropy loss; F s , F q are the features of the support set and the query set after passing through the self-feature extraction module respectively; A s , A q refer to the attention maps corresponding to the support set and the query set respectively; y is the sample label; sim is the cosine similarity; τ is the scalar temperature factor; λ is the hyperparameter used to balance the two losses; (5) Few-shot classification: According to the loss calculation result in step (4), classify the query set and output the final classification result; The few-shot classification step classifies the query set into the closest support set category by calculating the similarity between the query set and the support set.