A pulse neural network image classification method and device based on isomorphic model progressive training

CN121708357BActive Publication Date: 2026-09-18ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511764080.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-27
Publication Date
2026-09-18
Estimated Expiration
2045-11-27

AI Technical Summary

Technical Problem

然而,由于脉冲信号的离散性和不可微特性,SNN难以直接应用基于梯度的反向传播算法进行端到端训练,导致其训练过程收敛困难、优化不稳定,在相同网络架构下,SNN的训练精度显著低于传统人工神经网络(ANN)

Benefits of technology

[0013]The beneficial effects of this invention are as follows: This invention proposes a method and apparatus for image classification using a spiking neural network (SNN) based on progressive training of isomorphic models. It eliminates the need to initialize a small-capacity SNN from scratch. Instead, it efficiently transfers knowledge from a large-capacity artificial neural network to a small-capacity model through structure-preserving weight transformation, smoothly transitioning to the spiking neural network, significantly improving the convergence speed and final accuracy of the SNN. The structured weight transformation strategy employed strictly maintains the relative positional relationship and channel structure of the model weights, maximizing the retention of pre-trained knowledge. In particular, this method is applicable to any family of isomorphic models (such as RegNet) and image classification tasks, exhibiting good versatility and scalability, and improving convergence accuracy in image classification tasks. This invention effectively narrows the performance gap between SNNs and artificial neural networks (ANNs) without increasing model complexity or inference energy consumption, significantly improving the final accuracy of SNNs in image classification tasks, and providing an effective solution for the practical deployment of lightweight, high-precision spiking neural networks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121708357B_ABST
    Figure CN121708357B_ABST
Patent Text Reader

Abstract

The application discloses a kind of based on isomorphic model progressive training's pulse neural network image classification method and device.Firstly, a large capacity artificial neural network is trained as source model;Then the structural consistency of isomorphic model in the same model family is used, through the structural weight conversion of depth and width dimension, the source model weight is migrated to small capacity artificial neural network for initialization and fine tuning;Finally, the small capacity model is converted into the pulse neural network of the same structure and completes training.The relationship consistent interpolation and sequential mapping strategy used in the application effectively maintains the relative position relationship of weight between isomorphic models, maximizes the pre-training knowledge retention;And through the two-stage migration mechanism of progressive, realize the smooth transition of knowledge, provide high-quality initialization for pulse neural network.The application significantly improves the final accuracy of SNN in image classification task, provides an effective solution for the practical deployment of lightweight high-precision pulse neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of neural network and image classification technology, specifically relating to a spiking neural network image classification method and apparatus based on progressive training of isomorphic models. Background Technology

[0002] Spiking Neural Networks (SNNs), as an important carrier of neuromorphic computing, simulate the information transmission mechanism of biological neurons through discrete pulse signals. They possess advantages such as event-driven operation, low power consumption, and high temporal resolution, making them particularly suitable for neuromorphic hardware and edge intelligence scenarios. However, due to the discreteness and non-differentiability of pulse signals, SNNs are difficult to train end-to-end using gradient-based backpropagation algorithms. This leads to difficulties in convergence and unstable optimization during training, resulting in significantly lower training accuracy compared to traditional Artificial Neural Networks (ANNs) under the same network architecture. This performance gap severely restricts the practical deployment of SNNs in image classification tasks.

[0003] To improve SNN performance, existing methods often rely on alternative gradient training from scratch or complex temporal encoding strategies. However, these methods are sensitive to hyperparameters, have high training costs, and struggle to achieve ideal results on lightweight models. Especially when the target SNN has a small capacity, random initialization often fails to provide an effective starting point for optimization, further exacerbating accuracy loss. Although the isomorphic design of model families (such as RegNet and ResNet) has provided a structural foundation for cross-model knowledge transfer in recent years, existing SNN training frameworks have not fully utilized this characteristic. Therefore, there is an urgent need for a training method that can leverage the structural consistency between isomorphic models to achieve efficient initialization and progressive optimization, thereby overcoming the accuracy bottleneck of SNNs in image classification tasks. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention proposes a method and apparatus for image classification using spiking neural networks based on progressive training of isomorphic models. This method utilizes the structural consistency between isomorphic models within the same model family to transfer the weights of a trained large-capacity artificial neural network to a small-capacity artificial neural network for initialization and training through structured weight transformation. Furthermore, the small-capacity artificial neural network is converted into a structurally isomorphic spiking neural network for training, thereby improving the convergence accuracy of the final spiking neural network in image classification tasks.

[0005] In a first aspect, the present invention provides a spiking neural network image classification method based on progressive training of an isomorphic model, the method comprising the following steps: (1) Collect target image data, construct an image classification dataset, train a large-capacity artificial neural network as the source model, and obtain the weights of the source model; (2) Utilizing the structural consistency between isomorphic models, the weights of the source model are transferred to a small-capacity artificial neural network of the same model family through structured weight transformation for initialization and training and fine-tuning based on an image classification dataset; (3) The trained small-capacity artificial neural network is converted into a structurally isomorphic spiking neural network and trained on the same image classification dataset; (4) Image classification is achieved based on the trained spiking neural network to improve convergence accuracy.

[0006] Furthermore, the isomorphic model refers to neural networks belonging to the same model family, which have the same macroscopic architectural design and basic building blocks. Each model includes an input structure stem, a network body structure, and an output structure head connected in sequence; wherein: (1) The input structure includes a convolutional layer and a batch normalization layer; (2) The main structure of the network includes multiple computation stages, each of which is composed of several residual basic units stacked together. The residual basic unit adopts the residual connection structure of identity mapping and contains at least two convolutional layers, corresponding batch normalization layers and activation functions. (3) For image classification tasks, the output structure includes a global average pooling layer and a fully connected layer, and the number of output channels of the fully connected layer is equal to the number of image categories.

[0007] Furthermore, isomorphic models belong to the same model family, and family members differ only in model depth, model width, and activation function type. Their weights can be transferred between each other through structured weight transformation methods.

[0008] Furthermore, the structured weight transformation and structural isomorphism transformation are both achieved through transformations in the following three dimensions: (1) Depth-dimensional transformation: For each computation stage in the main network structure, based on the difference in the number of residual basic units contained in the source model and the target model in that stage, a sequential mapping relationship between units is established; when the depth of the target model is less than that of the source model, based on the number of units in the target model, starting from the first unit, the weights of the corresponding number of units in the source model are retained; when the depth of the target model is greater than that of the source model, all the weights of the source model are retained, and the weights of the last unit of the source model are copied to fill the redundant units; this strategy ensures that the relative positional relationship of the residual basic units in the stage is maintained, thereby maintaining the hierarchical feature extraction structure learned in the source model; (2) Width dimension transformation: For the weight tensors of convolutional layers, batch normalization layers and fully connected layers, the relational consistency interpolation method is used to perform multi-dimensional interpolation along the input channel and output channel dimensions, so as to maintain the relative channel relationship in the original weights to the greatest extent and retain the pre-trained knowledge contained in the weights; (3) Activation function adaptation: When converting an artificial neural network into a spiking neural network, the continuous activation function in the original artificial neuron is replaced with a spiking neuron model, which includes an integral-and-fire (IF) neuron. Since neither the continuous activation function nor the spiking neuron model contains learnable parameters, this replacement can be performed directly without changing the network weights.

[0009] Furthermore, the specific steps of the relational consistency interpolation method are as follows: For a one-dimensional input vector and target dimension Calculation quotient and remainder ; Indicates the dimension of the input vector, if The output will be ; Otherwise, repeat The intermediate vector is obtained once, and... Concatenate to the end of the intermediate vector to form a vector of length . The output vector; For a multidimensional weight tensor, apply the one-dimensional interpolation operation described above along the channel dimensions to be transformed.

[0010] Secondly, the present invention provides a spiking neural network image classification device based on isomorphic model progressive training, including a memory and one or more processors. The memory stores executable code, and when the processor executes the executable code, it implements the spiking neural network image classification method based on isomorphic model progressive training.

[0011] Thirdly, the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the aforementioned image classification method for a spiking neural network based on progressive training of an isomorphic model.

[0012] Fourthly, the present invention provides a computer program product, including a computer program, which, when executed by a processor, implements the aforementioned image classification method for spiking neural networks based on isomorphic model progressive training.

[0013] The beneficial effects of this invention are as follows: This invention proposes a method and apparatus for image classification using a spiking neural network (SNN) based on progressive training of isomorphic models. It eliminates the need to initialize a small-capacity SNN from scratch. Instead, it efficiently transfers knowledge from a large-capacity artificial neural network to a small-capacity model through structure-preserving weight transformation, smoothly transitioning to the spiking neural network, significantly improving the convergence speed and final accuracy of the SNN. The structured weight transformation strategy employed strictly maintains the relative positional relationship and channel structure of the model weights, maximizing the retention of pre-trained knowledge. In particular, this method is applicable to any family of isomorphic models (such as RegNet) and image classification tasks, exhibiting good versatility and scalability, and improving convergence accuracy in image classification tasks. This invention effectively narrows the performance gap between SNNs and artificial neural networks (ANNs) without increasing model complexity or inference energy consumption, significantly improving the final accuracy of SNNs in image classification tasks, and providing an effective solution for the practical deployment of lightweight, high-precision spiking neural networks. Attached Figure Description

[0014] Figure 1 This is a structural diagram of an isomorphic model provided in an embodiment of the present invention; Figure 2 The diagram shows the structure of a spiking neural network image classification device based on progressive training of an isomorphic model, as provided by the present invention. Detailed Implementation

[0015] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.

[0016] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0017] This invention proposes a spiking neural network image classification method based on progressive training of isomorphic models. This method utilizes the structural consistency between isomorphic models in the same model family. The weights of a large-capacity artificial neural network trained on an image classification dataset are transferred to a small-capacity artificial neural network for initialization through structured weight transformation. The small-capacity artificial neural network is then trained on the same image classification data. Furthermore, the small-capacity artificial neural network is converted into a structurally isomorphic spiking neural network and trained on the same image classification data, thereby improving the convergence accuracy of the final spiking neural network in the image classification task.

[0018] The homogeneous model refers to a neural network belonging to the same model family, which has the same macroscopic architectural design and basic building blocks, such as... Figure 1 As shown. Each model consists of an input structure (stem), a network body structure, and an output structure (head) connected in sequence; where: (1) The input structure includes a convolutional layer and a batch normalization layer; (2) The main structure of the network includes multiple computation stages, each of which is composed of several residual basic units stacked together. The residual basic unit adopts the residual connection structure of identity mapping and contains at least two convolutional layers, corresponding batch normalization layers and activation functions. (3) For image classification tasks, the output structure includes a global average pooling layer and a fully connected layer, and the number of output channels of the fully connected layer is equal to the number of image categories.

[0019] Different isomorphic models differ mainly in model depth, model width, and activation function type, and their weights can be transferred to each other through structured weight transformation methods.

[0020] The structured weight transformation method includes transformations in the following three dimensions: (1) Depth Dimension Transformation: For each computation stage in the main network structure, an order mapping relationship between units is established based on the difference in the number of residual basic units contained in the source model and the target model in that stage. When the depth of the target model is less than that of the source model, the weights of the first N units of the source model are retained. When the depth of the target model is greater than that of the source model, the weights of the last unit of the source model are copied to fill the redundant units. This strategy ensures that the relative positional relationship of the residual basic units in the stage is maintained, thereby maintaining the hierarchical feature extraction structure learned in the source model.

[0021] (2) Width dimension transformation: For the weight tensors of convolutional layers, batch normalization layers and fully connected layers, the relational consistency interpolation method is used to perform multidimensional interpolation along the input channel and output channel dimensions, so as to maintain the relative channel relationship in the original weights to the greatest extent and retain the pre-trained knowledge contained in the weights.

[0022] (3) Activation function adaptation: When converting an artificial neural network into a spiking neural network, the continuous activation function in the original artificial neuron is replaced with a spiking neuron model, which includes, but is not limited to, an integral-and-fire (IF) neuron. Since neither the continuous activation function nor the spiking neuron model contains learnable parameters, this replacement can be performed directly without changing the network weights.

[0023] The specific steps of the relation-consistent interpolation method are as follows: For a one-dimensional input vector and target dimension Calculation quotient and remainder ; Indicates the dimension of the input vector; if The output will be ; Otherwise, repeat The intermediate vector is obtained once, and... Concatenate to the end of the intermediate vector to form a vector of length . The output vector; For a multidimensional weight tensor, apply the one-dimensional interpolation operation described above along the channel dimensions to be transformed.

[0024] The progressive training process includes the following steps: (1) Train a large-capacity artificial neural network on an image classification dataset as the source model; (2) Select a small-capacity artificial neural network from the same model family as an intermediate model, initialize it by transforming the weights of the source model through the depth and width dimensions, and fine-tune it on the same dataset; (3) The activation function is adapted to the fine-tuned intermediate model, and the continuous activation function is replaced with spiking neurons to construct the target spiking neural network; (4) On the same image classification dataset, the target spiking neural network is trained using the gradient substitution backpropagation algorithm to obtain the final model.

[0025] To demonstrate the advancements of the proposed method, the proposed spiking neural network image classification method based on isomorphic model progressive training is compared with different spiking neural network initialization methods on multiple image classification datasets. The target spiking neural network uses the parameter configuration of RegNet200MF, and the source large-capacity model uses the parameter configuration of RegNet600MF.

[0026] CIFAR100 is a commonly used dataset for image classification tasks, containing 60,000 RGB images of various sizes. These images cover diverse categories, including animals, fruits, people, vehicles, and flowers, with a total of 100 categories. TinyImageNet is a subset of the ImageNet dataset, containing 100,000 training images and 10,000 test images, with 200 categories. CUB200 is a flower image classification dataset, containing 5,994 training images and 5,784 test images, with 200 categories. Caltech 256 contains 25,611 training images and 4,996 test images, with 257 categories.

[0027] Table 1 Table 1 compares the accuracy of the proposed spiking neural network image classification method based on isomorphic model progressive training with other spiking neural network initialization methods. It can be seen that, with identical model structure and training settings, the proposed method significantly improves image classification accuracy across various datasets.

[0028] Corresponding to the aforementioned embodiment of a spiking neural network image classification method based on isomorphic model progressive training, the present invention also provides an embodiment of a spiking neural network image classification device based on isomorphic model progressive training.

[0029] See Figure 2 The present invention provides a spiking neural network image classification device based on isomorphic model progressive training, comprising a memory and one or more processors. The memory stores executable code, and when the processor executes the executable code, it is used to implement a spiking neural network image classification method based on isomorphic model progressive training as described in the above embodiment.

[0030] The embodiment of the spiking neural network image classification device based on isomorphic model progressive training provided by this invention can be applied to any device with data processing capabilities, such as a computer. The device embodiment can be implemented in software, hardware, or a combination of both. Taking software implementation as an example, as a logical device, it is formed by the processor of any data processing device loading the corresponding computer program instructions from non-volatile memory into memory for execution. From a hardware perspective, such as... Figure 2 The diagram shown is a hardware structure diagram of any device with data processing capabilities, which includes a spiking neural network image classification device based on isomorphic model progressive training provided by the present invention. Except for... Figure 2 In addition to the processor, memory, network interface, and non-volatile memory shown, any data processing device in the embodiment may also include other hardware depending on the actual function of the data processing device, which will not be described in detail here.

[0031] The specific implementation process of the functions and roles of each unit in the above device can be found in the implementation process of the corresponding steps in the above method, and will not be repeated here.

[0032] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to in the description of the method embodiments. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of the present invention according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0033] This invention also provides a computer-readable storage medium storing a program thereon, which, when executed by a processor, implements a spiking neural network image classification method based on isomorphic model progressive training as described in the above embodiments.

[0034] The computer-readable storage medium can be an internal storage unit of any data processing device described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium can also be an external storage device of any data processing device, such as a plug-in hard disk, smart media card (SMC), SD card, flash card, etc., equipped on the device. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of any data processing device. The computer-readable storage medium is used to store the computer program and other programs and data required by the data processing device, and can also be used to temporarily store data that has been output or will be output.

[0035] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the aforementioned image classification method for spiking neural networks based on isomorphic model progressive training.

[0036] The above description is merely a preferred embodiment of the present invention. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make many possible variations and modifications to the technical solutions of the present invention using the methods and techniques disclosed above, or modify them into equivalent embodiments with equivalent changes, without departing from the scope of the technical solutions of the present invention. Therefore, any simple modifications, equivalent changes, and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention shall still fall within the protection scope of the technical solutions of the present invention.

Claims

1. A spiking neural network image classification method based on isomorphic model progressive training, characterized in that, The method includes the following steps: (1) Collect target image data, construct an image classification dataset, train a large-capacity artificial neural network as the source model, and obtain the weights of the source model; (2) Utilizing the structural consistency between isomorphic models, the weights of the source model are transferred to a small-capacity artificial neural network of the same model family for initialization through structured weight transformation, and then trained and fine-tuned based on an image classification dataset; the isomorphic models belong to the same model family, and the family members differ only in model depth, model width, and activation function type, and their weights can be transferred to each other through the structured weight transformation method; both the structured weight transformation and the structural isomorphic transformation are achieved through the following three dimensions of transformation: 1) Depth-dimensional transformation: For each computation stage in the main network structure, establish a sequential mapping relationship between units based on the difference in the number of residual basic units contained in the source model and the target model at that stage; when the depth of the target model is less than that of the source model, based on the number of units in the target model, starting from the first unit, retain the weights of the corresponding number of units in the source model; when the depth of the target model is greater than that of the source model, retain all the weights of the source model, and copy the weights of the last unit of the source model to fill the redundant units; 2) Width dimension transformation: For the weight tensors of convolutional layers, batch normalization layers and fully connected layers, a relational consistent interpolation method is used to perform multidimensional interpolation along the input and output channel dimensions; 3) Activation function adaptation: When converting an artificial neural network into a spiking neural network, the continuous activation function in the original artificial neuron is replaced with a spiking neuron model, which includes an integral-fire neuron. Since neither the continuous activation function nor the spiking neuron model contains learnable parameters, this replacement can be performed directly without changing the network weights. (3) The trained small-capacity artificial neural network is converted into a structurally isomorphic spiking neural network and trained on the same image classification dataset; (4) Image classification is achieved based on the trained spiking neural network to improve convergence accuracy.

2. The image classification method for spiking neural networks based on progressive training of isomorphic models according to claim 1, characterized in that, The isomorphic model refers to a neural network belonging to the same model family. They have the same macroscopic architecture design and basic building blocks. Each model includes an input structure stem, a network body structure, and an output structure head connected in sequence; wherein: (1) The input structure includes a convolutional layer and a batch normalization layer; (2) The main structure of the network includes multiple computation stages, each of which is composed of several residual basic units stacked together. The residual basic unit adopts the residual connection structure of identity mapping and contains at least two convolutional layers, corresponding batch normalization layers and activation functions. (3) For image classification tasks, the output structure includes a global average pooling layer and a fully connected layer, and the number of output channels of the fully connected layer is equal to the number of image categories.

3. The image classification method for spiking neural networks based on progressive training of isomorphic models according to claim 1, characterized in that, The specific steps of the relation-consistent interpolation method are as follows: For a one-dimensional input vector and target dimension Calculation quotient and remainder ; Indicates the dimension of the input vector, if The output will be ; Otherwise, repeat The intermediate vector is obtained once, and... Concatenate to the end of the intermediate vector to form a vector of length [length missing]. The output vector; For a multidimensional weight tensor, apply the one-dimensional interpolation operation described above along the channel dimensions to be transformed.

4. A spiking neural network image classification device based on isomorphic model progressive training, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that... When the processor executes the executable code, it implements a spiking neural network image classification method based on progressive training of an isomorphic model as described in any one of claims 1-3.

5. A computer-readable storage medium having a program stored thereon, characterized in that, When the program is executed by the processor, it implements a spiking neural network image classification method based on progressive training of an isomorphic model as described in any one of claims 1-3.

6. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the spiking neural network image classification method based on progressive training of isomorphic models as described in any one of claims 1-3.

Citation Information

Patent Citations

  • Width pulse neural network-based image classification model training method and device

    CN118968151A

  • Multi-stage pulse neural network training method and device based on knowledge distillation

    CN119129702A