Knowledge distillation method based on interactive combination of ANN and SNN

Through the interactive joint training and knowledge distillation between ANN and SNN, the problem of high complexity of ANN overfitting and SNN training is solved, and the performance and computing efficiency of SNN in complex timing tasks is improved. It is suitable for low-power devices, especially in tasks such as video analysis and speech recognition.

CN120258086APending Publication Date: 2025-07-04SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510291007.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

ANNs are prone to overfitting when processing complex modes and large-scale data and lack the processing capability of timing dynamics. SNNs have high training complexity and are difficult to widely use in general computing platforms. In addition, traditional SNNs are difficult to efficiently integrate historical and current features in dynamic scenarios, resulting in insufficient adaptability to rapidly changing input data, serious resource waste, and it is difficult to meet the real-time needs of low-power devices.

Method used

By constructing an interactive joint network architecture with ANN as the teacher model and SNN as the student model, the knowledge distillation method is used to minimize the differences and losses between the teacher model and the student model output, and feature fusion is combined with attention mechanism and historical time step information to optimize the training process.

Benefits of technology

It improves the performance of SNN in handling complex timing tasks, solves the time step alignment problem of ANN and SNN, enhances the robustness and real-timeness of the model for dynamic scenarios, reduces computing resource consumption, and is suitable for embedded devices and edge computing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258086A_ABST
    Figure CN120258086A_ABST
Patent Text Reader

Abstract

The invention discloses a knowledge distillation method based on interactive combination of ANN and SNN, and the method comprises the following steps: S1, obtaining a visual image, carrying out the preprocessing of the visual image, and inputting the processed data into a neural network; s2, the input data is processed through an ANN; s3, processing the input data through the SNN; and S4, constructing an interactive joint network architecture with the ANN as a teacher model and the SNN as a student model, minimizing the difference and loss between the output of the teacher model and the output of the student model through knowledge distillation, and guiding the training of the student model. The invention further discloses a knowledge distillation system based on the interactive joint of the ANN and the SNN. Comprising a data acquisition module, a feature extraction module, a knowledge distillation module, an interactive learning module and a classification evaluation module. According to the method, through joint training and knowledge distillation of the ANN and the SNN, the performance of the SNN is improved by utilizing the high-level feature extraction capability of the ANN and the time sequence processing advantages of the SNN, and the calculation efficiency and the energy efficiency are optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of deep learning, and in particular to a knowledge distillation method based on the interaction and combination of ANN and SNN. Background Art

[0002] With the rapid development of artificial intelligence and deep learning technologies, artificial neural networks (ANNs) and spiking neural networks (SNNs) have been widely used in various intelligent systems. ANNs have received extensive attention due to their efficient performance in tasks such as image recognition and natural language processing. However, the training process of ANNs usually requires a large amount of computing resources, and the models are difficult to process temporal continuity or sequential information. In contrast, SNNs simulate the working principle of biological nervous systems, can efficiently process sequential data, and have obvious advantages in terms of computational efficiency and energy efficiency, especially suitable for embedded devices and low-power environments.

[0003] Although ANNs and SNNs each have their own unique advantages, they also have some limitations. ANNs are prone to overfitting when dealing with complex patterns and large-scale data and lack the ability to process temporal dynamics, while the training complexity of SNNs is high and the hardware requirements are relatively high, making it difficult to be widely applied to general computing platforms. Therefore, how to make full use of the complementary advantages of ANNs and SNNs and achieve information sharing between the two through an effective knowledge distillation method has become a hot research issue.

[0004] As a model compression technology, knowledge distillation improves the generalization ability of a small model (student model) by transferring the knowledge of a large and complex model (teacher model). In the joint training of ANNs and SNNs, the application of knowledge distillation can help the SNN network learn more high-level features from the ANN while retaining its advantage in temporal processing.

[0005] Therefore, the technical personnel in this field are committed to developing a knowledge distillation method based on the interaction and combination of ANNs and SNNs. Summary of the Invention

[0006] In view of the above-mentioned defects of the prior art, the technical problems to be solved by the present invention are as follows: ANN is prone to overfitting when dealing with complex patterns and large-scale data and lacks the ability to process temporal dynamics, while the training complexity of SNN is high and the hardware requirements are relatively high, making it difficult to be widely applied to general computing platforms; there is a problem of time step mismatch when distilling ANN to SNN, and ANN cannot guide the temporal inference process of SNN in terms of time steps; traditional SNN is difficult to efficiently fuse historical and current features in dynamic scenarios, resulting in insufficient adaptability to rapidly changing input data; the traditional inference process requires fixed time step calculation, leading to waste of resources and difficulty in meeting the real-time requirements of low-power devices.

[0007] To achieve the above object, the present invention provides a knowledge distillation method based on the interaction and combination of ANN and SNN, and the method includes the following steps:

[0008] S1: Obtain a visual image, preprocess the visual image, and input the processed data into a neural network;

[0009] S2: Process the input data through ANN;

[0010] S3: Process the input data through SNN;

[0011] S4: Construct an interactive combined network architecture with ANN as the teacher model and SNN as the student model, and minimize the difference and loss between the outputs of the teacher model and the student model through knowledge distillation to guide the training of the student model;

[0012] Further, the visual image in step S1 can be collected in real time through an event-driven vision sensor, a camera or a lidar;

[0013] Further, the preprocessing in step S1 includes:

[0014] Perform bilinear interpolation scaling on the visual image I orig to reduce its size to the target size I small , which is represented by the formula:

[0015] I small (x, y) = I orig (r * x, r * y),

[0016] where r is the reduction ratio, and x, y are the pixels of the image;

[0017] Further, the preprocessing in step S1 further includes: performing gray-scale enhancement and edge-preserving filtering operations on the image of the target size;

[0018] Further, in step S2, it is processed by an ANN to generate a high-level feature representation, which is expressed by the formula:

[0019] H ann = f ann (X)

[0020] where H ann is the feature representation output by the ANN, f snn is the forward propagation function of the ANN, and X is the corresponding input;

[0021] Further, the processing by the SNN in step S3 includes gradually updating the membrane potential through time steps, and the final temporal feature representation is:

[0022] H snn = f snn (X(t))

[0023] where H snn is the feature representation output by the SNN, f snn is the forward propagation function of the SNN, and X is the corresponding input;

[0024] The specific formula of the LIF activation function is expressed as:

[0025] V[t] = H[t](1 - S[t]) + V reset S[t],

[0026]

[0027] where V[t] represents the membrane potential at time step t, S[t] represents the spike firing state at time step t, V threshold represents the threshold voltage, V reset represents the reset voltage, H[t] represents the intermediate variable at time step t, τ is the time constant, and the membrane potential V[t] is gradually updated according to the input X[t]. When V[t] exceeds the threshold voltage V threshold , a spike is fired (i.e., S[t] = 1), generating a binary feature map S[t]; X[t] is the feature map at time step t;

[0028] Further, the difference and loss in step S4 are expressed by the formula:

[0029]

[0030] where L KD is the knowledge distillation loss, which measures the difference between the outputs of the student model and the teacher model, T is the number of time steps, γ t is the weighting coefficient at time step t, and KL(y t+1 ||y t) is the KL divergence, which compares the differences in the predicted probability distributions between time steps t and t+1. and is the predicted probability of the i-th class at time step t+1 and t; L CE (y T , y true ) is the cross-entropy loss, which measures the difference between the class probability distribution predicted by the model and the true label; for the outputs {y1, y3, y3,..., y T} at T time steps, CE represents the final cross-entropy loss, and γ t is the weight parameter that controls the influence degree of the guidance loss between different time steps. y t+1 and y t represent the output probability distributions at the (t+1)-th and t-th time steps respectively. represents the probability value of this distribution in the i-th class;

[0031] Using the cross-entropy loss, y i represents the value of the i-th class in the true label, which is set to a one-hot encoded value, and p i is the predicted probability of the i-th class output by the model;

[0032] Using the L2 norm, calculate the squared difference between p1 and C(p1). C represents the convolution operation on p1 to measure the difference between feature maps;

[0033] Furthermore, jointly train the teacher model and the student model. At each time step, identify the key regions through the attention mechanism and perform cross-time-step feature fusion by combining the information of historical time steps, which is represented by the formula:

[0034] H fusion (t) = f fusion (H snn (t), H ann (t), Hf usion (t - 1))

[0035] where H fusion (t) is the fused feature at the current time step, and f fusion is the feature fusion function;

[0036] The present invention also provides a knowledge distillation system based on the interaction and combination of ANN and SNN, including a data acquisition module, a feature extraction module, a knowledge distillation module, an interaction learning module, and a classification evaluation module:

[0037] The data acquisition module is used to obtain data in real time and preprocess the data;

[0038] The feature extraction module extracts high-dimensional features from the preprocessed data through a neural network to generate an abstract representation characterizing the key features;

[0039] The knowledge distillation module transfers the knowledge of the teacher model to the student model and guides the training of the student model by calculating the difference between the outputs of the teacher model and the student model;

[0040] The interactive learning module enables the teacher model and the student model to transfer information to each other through joint training to achieve collaborative learning between the two;

[0041] The classification and evaluation module makes a classification decision on the output features after feature extraction and knowledge distillation, and feeds back to other modules through accuracy evaluation to further optimize the system performance;

[0042] Further, the difference includes using KL divergence to measure the difference in the output distributions of the teacher model and the student model.

[0043] Through the joint training and knowledge distillation of ANN and SNN, the present invention enables SNN to extract high-level abstract features from ANN, which are crucial for processing complex temporal tasks (such as video analysis, speech recognition, etc.). Compared with traditional single SNN models, joint training enables SNN to perform better in the face of dynamic scenarios and rapidly changing environments; solves the time step alignment problem between ANN and SNN and improves the coherence of temporal reasoning; accelerates model convergence through dynamic weight adjustment and reduces training time; enhances the robustness of the model to dynamic scenarios (such as fast-moving targets, sudden noises); optimizes current predictions through historical information and improves the real-time performance of temporal tasks (such as autonomous driving, video surveillance); reduces redundant calculations, lowers inference energy consumption, and is applicable to embedded devices and edge computing scenarios; and improves the response speed of real-time tasks while ensuring accuracy. Brief Description of the Drawings

[0044] Figure 1 It is a schematic diagram of the architecture of the knowledge distillation method based on the interactive combination of ANN and SNN of the present invention;

[0045] Figure 2 It is a schematic diagram of the working process of the knowledge distillation method based on the interactive combination of ANN and SNN of the present invention;

[0046] Figure 3 It is a schematic diagram of the loss architecture of the teacher model and the student model of the present invention;

[0047] Figure 4 It is a schematic diagram of the structure of the knowledge distillation system based on the interactive combination of ANN and SNN of the present invention. Detailed Embodiments

[0048] The following describes several preferred embodiments of the present invention with reference to the accompanying drawings of the specification to make its technical content clearer and easier to understand. The present invention can be embodied in many different forms of embodiments, and the protection scope of the present invention is not limited to the embodiments mentioned in the text.

[0049] In the accompanying drawings, components with the same structure are denoted by the same numerical reference signs, and components with similar structures or functions everywhere are denoted by similar numerical reference signs. The size and thickness of each component shown in the accompanying drawings are arbitrarily shown, and the present invention does not limit the size and thickness of each component. To make the illustration clearer, the thickness of some components in the accompanying drawings is appropriately exaggerated.

[0050] To solve the efficiency problem in transfer learning between deep learning models in the prior art, the present invention proposes a knowledge distillation method based on the interaction and combination of ANN and SNN. This method improves the performance of SNN and optimizes the computational efficiency and energy efficiency by jointly training ANN and SNN, taking advantage of the high-level feature extraction ability of ANN and the temporal processing advantage of SNN. Specifically, the present invention adopts the knowledge distillation technology to transfer the knowledge obtained from the training of ANN to SNN, so that SNN can better learn the important features in complex temporal data and improve its performance in temporal tasks. The schematic diagram of the architecture of the knowledge distillation method of the present invention is shown in FIG. 1, and this method includes the following steps:

[0051] S1: Obtain a visual image, preprocess the visual image, and input the processed data into a neural network;

[0052] S2: Process the input data through ANN;

[0053] S3: Process the input data through SNN;

[0054] S4: Construct an interaction and combination network architecture with ANN as the teacher model and SNN as the student model, and minimize the difference and loss between the outputs of the teacher model and the student model through knowledge distillation to guide the training of the student model.

[0055] In a more specific embodiment, as Figure 2 shown in the schematic diagram of the working process of the knowledge distillation method based on the interaction and combination of ANN and SNN of the present invention, the working process includes:

[0056] Obtain a visual image and preprocess the visual image;

[0057] In this embodiment, the visual image can be used to collect visual image data in real time through devices such as an event-driven vision sensor camera;

[0058] In this embodiment, for the visual image I origPerform bilinear interpolation scaling to reduce its size to the target size I small , this process not only helps to reduce the interference of irrelevant information, but also maintains the overall structure of the image, enabling the features of the key regions to be retained. It is expressed by the formula:

[0059] I small (x, y) = I orig (r * x, r * y),

[0060] where r is the scaling ratio (e.g., r = 0.5 means reducing to half the size), and x, y are the pixels of the image.

[0061] To ensure the visual features of the downsized image are clear, further preprocessing is performed on the downsized image, including operations such as gray-scale enhancement and edge-preserving filtering, to maintain the clarity of the image.

[0062] Furthermore, the input data is processed through an artificial neural network (ANN) to generate a high-level feature representation, which is expressed by the formula:

[0063] H ann = f ann (X)

[0064] where H ann is the feature representation output by the ANN, f ann is the forward propagation function of the ANN, and X is the corresponding input. In this embodiment, multi-layer convolution operations are performed on the preprocessed image, and channel feature maps are generated corresponding to each layer of convolution.

[0065] Similarly, the input data is also processed through a spiking neural network (SNN). The SNN simulates the temporal processing mechanism of the biological nervous system and gradually updates the membrane potential through time steps to finally obtain a temporal feature representation:

[0066] H snn = f snn (X(t))

[0067] where H snn is the feature representation output by the SNN, f snn is the forward propagation function of the SNN, and X is the corresponding input. The specific formula of the LIF activation function is expressed as:

[0068] V[t] = H[t](1 - S[t]) + V reset S[t],

[0069]

[0070] where V[t] represents the membrane potential at time step t, S[t] represents the spike emission state at time step t, and Vthreshold Denotes the threshold voltage, V reset Denotes the reset voltage, H[t] represents the intermediate variable at time step t, τ is the time constant, and the membrane potential V[t] is gradually updated according to the input X[t]. When V[t] exceeds the threshold voltage V threshold , a pulse is emitted (i.e., S[t]=1), generating a binary feature map S[t]; X[t] is the feature map at time step t.

[0071] After ANN and SNN feature extraction, a knowledge distillation process is performed. Knowledge distillation enables the SNN to learn the feature representation of the ANN by transferring the high-level features of the ANN to the SNN. The goal of knowledge distillation is to minimize the loss between the output of the SNN and the output of the ANN, guiding the training of the student model.

[0072] As Figure 3 shown in the schematic diagram of the loss architecture of the teacher model and the student model of the embodiment of the present invention, the loss function includes:

[0073]

[0074] Among them, L KD is the knowledge distillation loss, measuring the difference between the output of the student model and the output of the teacher model, T is the number of time steps, γt is the weighting coefficient at time step t, KL(y t+1 ||y t ) is the KL divergence, comparing the difference in the predicted probability distributions between time steps t and t+1, and are the predicted probabilities of the i-th category at time steps t+1 and t; L CE (y T ,y true ) is the cross-entropy loss, used to measure the difference between the category probability distribution predicted by the model and the true label.

[0075] For the output {y1,y3,y3,...,y T} of T time steps, CE represents the final cross-entropy loss, γ t is the weight parameter, used to control the influence degree of the guiding loss between different time steps, y t+1 and y t represent the output probability distributions at the (t+1)-th and t-th time steps respectively, represents the probability value of this distribution in the i-th category.

[0076] By minimizing the KL divergence, y t can be guided to be closer to the distribution of y t+1 , thereby gradually approaching the final prediction result at the early time steps.

[0077] Using cross - entropy loss, y i represents the value of the i - th class in the true label, set as a one - hot encoded value, p i is the predicted probability of the i - th class output by the model; Using the L2 norm, calculate the squared difference between p1 and C(p1), where C represents the convolution operation on p1 to measure the difference between feature maps.

[0078] The time - step guidance mechanism based on knowledge distillation aims to accelerate the model inference process. By gradually guiding the output distribution of the previous time - steps to approach that of the subsequent time - steps, it can quickly obtain the features required for the final result. This mechanism introduces the KL divergence in the loss function and uses the output of the subsequent time - steps as the guidance signal, enabling the previous time - steps to gradually learn the key information of the subsequent time - steps and optimize the inference process. Specifically, the model first quickly extracts global features at a low resolution and dynamically determines whether to further refine the features based on the confidence during the inference process. When the preliminary result meets a certain confidence level, the inference can end early, avoiding unnecessary calculations and improving efficiency.

[0079] During the inference stage, the model determines the similarity between the current time - step and the final time - step through an early - stopping condition, thereby dynamically deciding whether to end the inference early. This strategy reduces unnecessary consumption of computing resources and further improves the overall inference efficiency. Therefore, the time - step guidance mechanism based on knowledge distillation not only accelerates the model inference, improves the accuracy, but also optimizes the use of computing resources, especially performing outstandingly in real - time inference tasks.

[0080] Through the joint training of ANN and SNN, feature fusion is performed within multiple time - steps to ensure that the SNN can use historical information for more accurate prediction. At each time - step, the attention mechanism is used to identify key regions and cross - time - step feature fusion is carried out by combining the information of historical time - steps:

[0081] H fusion (t)=f fusion (H snn (t), H ann (t), H fusion (t - 1))

[0082] where H fusion (t) is the fused feature at the current time - step, and f fusion is the feature fusion function.

[0083] The knowledge distillation method based on the interaction and combination of ANN and SNN proposed by the present invention can provide efficient and stable support for the training and optimization of large models, significantly improving the performance and efficiency of the models when processing large-scale data. By combining the artificial neural network (ANN) with the spiking neural network (SNN) and adopting the knowledge distillation technology, it can effectively improve the performance of the SNN in processing complex time-series data tasks, while reducing the computational complexity and resource consumption, and is particularly suitable for the training and inference of large-scale models.

[0084] As Figure 4 shown, it is a schematic structural diagram of the knowledge distillation system based on the interaction and combination of ANN and SNN of the present invention. The system includes a data acquisition module, a feature extraction module, a knowledge distillation module, an interactive learning module, and a classification and evaluation module, where:

[0085] The data acquisition module is responsible for real-time acquisition of data from devices such as event-driven vision sensors, cameras, lidar, etc., and performs preprocessing to adapt to the subsequent feature extraction process. Through operations such as multi-scale sampling, image size adjustment, and denoising, the quality of the input data is ensured, and the processed data is input into the neural network. Its role is to ensure the real-time performance of the system and the effectiveness of the data, and provide high-quality input data for the subsequent deep learning models;

[0086] The feature extraction module uses methods such as convolutional neural network (CNN), artificial neural network (ANN), spiking neural network (SNN), etc. to extract high-dimensional features from the input data and generate feature representations suitable for subsequent analysis. The image is processed through standard convolutional layers, and activation functions such as ReLU introduce non-linear features, and the pooling layer reduces the computational complexity. The role of this module is to convert the original data into an abstract representation that can better characterize the key features, enabling the subsequent knowledge distillation and classification modules to perform efficient learning and inference based on these features;

[0087] The knowledge distillation module transfers the knowledge of the teacher model (ANN) to the student model (SNN). By calculating the difference between the outputs of the teacher model and the student model, especially using the KL divergence to measure the difference in their output distributions, the training of the student model is guided. The role of this module is to help the SNN learn more high-level feature representations, improve its performance in complex tasks, and at the same time effectively reduce the computational resource consumption to meet the requirements of low-power devices;

[0088] The interactive learning module realizes the collaborative learning between ANN and SNN, enabling the two models to exchange information through joint training. The high-level features extracted by ANN provide rich learning information for SNN, while SNN further optimizes the output of ANN based on temporal information. The role of this module is to promote the complementarity between models, enhance the temporal learning ability of SNN, and improve the adaptability and accuracy of the model in dynamic environments.

[0089] The classification and evaluation module makes classification decisions on the output features after feature extraction and knowledge distillation, and evaluates the accuracy of the classification results. Through traditional classification algorithms such as fully connected layers or support vector machines, the extracted features are used for final classification prediction. Its role is to make the final decision of the system, and feedback the accuracy evaluation to other modules to further optimize the system performance and ensure fast and accurate classification and target recognition in complex dynamic environments.

[0090] The present invention has significant practical value and industrialization potential in the field of artificial intelligence. First, by integrating the complementary advantages of artificial neural network (ANN) and spiking neural network (SNN), it solves the problems of overfitting of traditional models in static data and insufficient adaptability in temporal tasks. ANN is responsible for extracting spatial features of high-complexity data such as images and voices, while SNN efficiently processes temporal dynamic information based on biological neural mechanisms. The joint training and knowledge distillation technologies of the two significantly improve the accuracy and robustness of the model in tasks such as video analysis and speech recognition. Second, in the video action recognition task, it demonstrates the collaborative optimization of high accuracy and low energy consumption compared with traditional solutions.

[0091] The present invention supports low-cost deployment. At the software level, it provides seamless interfaces with frameworks such as TensorFlow and PyTorch. Through knowledge distillation technology, it reduces the demand for training data and shortens the training time. The dynamic inference mechanism reduces the dependence on the cloud and lowers the industrialization technical threshold. In the field of intelligent security, it can detect abnormal behaviors in video streams in real time and improve the monitoring efficiency in public places. In the field of autonomous driving, it can efficiently analyze multi-sensor temporal data and enhance the ability to respond to sudden road conditions. In the field of industrial Internet of Things, it supports the real-time analysis and fault warning of temporal signals such as equipment vibration and temperature. In addition, in edge devices such as drones and smart homes, the present invention can achieve low-latency voice interaction and image classification. The diversity and technical adaptability of these application scenarios reflect the market value and social benefits of the present invention, providing technical support for the industrial transformation of artificial intelligence technology.

[0092] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative efforts. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the protection scope determined by the claims.

Claims

1. A knowledge distillation method based on the interaction and combination of ANN and SNN, characterized in that, The method includes the following steps: S1: Obtain a visual image, preprocess the visual image, and input the processed data into a neural network; S2: Process the input data through an ANN; S3: Process the input data through an SNN; S4: Construct an interactive joint network architecture with the ANN as the teacher model and the SNN as the student model, and minimize the difference and loss between the outputs of the teacher model and the student model through knowledge distillation to guide the training of the student model.

2. The knowledge distillation method based on the interaction and combination of ANN and SNN according to claim 1, wherein In step S1, the visual image can be collected in real time by an event-driven vision sensor, a camera, or a lidar.

3. The knowledge distillation method based on the interaction and combination of ANN and SNN according to claim 2, wherein, The preprocessing in step S1 includes: Perform bilinear interpolation scaling on the visual image I orig to reduce its size to the target size I small , which is expressed by the formula: I small (x, y) = I orig (r * x, r * y), where r is the scaling ratio, and x and y are the pixels of the image.

4. The knowledge distillation method based on the interaction and combination of ANN and SNN according to claim 3, wherein, The preprocessing in step S1 further includes: performing gray-scale enhancement and edge-preserving filtering operations on the image of the target size.

5. The knowledge distillation method based on the interaction and combination of ANN and SNN according to claim 4, wherein In step S2, processing through the ANN is performed to generate a high-level feature representation, which is expressed by the formula: H ann = f ann (X) Among them, H ann is the feature representation output by the ANN, and f ann is the forward propagation function of the ANN, and X is the corresponding input.

6. The knowledge distillation method based on the interaction and combination of ANN and SNN according to claim 5, wherein The processing through the SNN in step S3 includes gradually updating the membrane potential through time steps, and the final temporal feature representation is: H snn = f snn (X(t)) Among them, H snn is the feature representation output by the SNN, and f snn is the forward propagation function of the SNN, and X is the corresponding input; The specific LIF activation function formula is expressed as: V[t] = H[t](1 - S[t]) + V reset S[t], Among them, V[t] represents the membrane potential at time step t, S[t] represents the spike firing state at time step t, V threshold represents the threshold voltage, V reset represents the reset voltage, H[t] represents the intermediate variable at time step t, τ is the time constant, and the membrane potential V[t] is gradually updated according to the input X[t]. When V[t] exceeds the threshold voltage V threshold , a spike is fired (i.e., S[t]=1), generating a binarized feature map S[t]; X[t] is the feature map at time step t.

7. The knowledge distillation method based on the interaction and combination of ANN and SNN according to claim 6, characterized in that, The difference and loss in step S4 are expressed by the formula: Among them, L KD is the knowledge distillation loss, which measures the difference between the outputs of the student model and the teacher model. T is the number of time steps, and γ t is the weighting coefficient at time step t. KL(y t+1 ||y t ) is the KL divergence, which compares the difference in the predicted probability distributions between time steps t and t + 1. and are the predicted probabilities for the i-th category at time steps t + 1 and t; L CE (y T , y true ) is the cross-entropy loss, which measures the difference between the categorical probability distribution predicted by the model and the true labels; for the outputs {y1, y2, y3,..., y T} at T time steps, CE represents the final cross-entropy loss, and γ t is the weight parameter that controls the influence degree of the guidance loss between different time steps. y t+1 and y t represent the output probability distributions at the (t + 1)-th and t-th time steps respectively. represents the probability value of this distribution for the i-th category. Using cross-entropy loss, y i represents the value of the i-th class in the true label, set as a one-hot encoded value, p i is the predicted probability of the i-th class output by the model; Using the L2 norm, calculate the squared difference between p1 and C(p1), where C represents the convolution operation on p1, to measure the difference between feature maps.

8. The knowledge distillation method based on the interaction and combination of ANN and SNN according to claim 7, wherein The teacher model and the student model are jointly trained. At each time step, the key region is identified through an attention mechanism, and feature fusion across time steps is performed by combining the information of historical time steps, which is expressed by the formula: H fusion h(t) = f fusion (H snn (t), H ann (t), H fusion (t - 1)) Among them, H fusion (t) is the fused feature at the current time step, and f fusion is the feature fusion function.

9. A knowledge distillation system based on the interaction between an ANN and an SNN, including a data acquisition module, a feature extraction module, a knowledge distillation module, an interactive learning module, and a classification and evaluation module, characterized in that the data acquisition module is used to obtain data in real time and preprocess the data; the feature extraction module extracts high-dimensional features from the preprocessed data through a neural network to generate an abstract representation characterizing the key features; the knowledge distillation module transfers the knowledge of the teacher model to the student model and guides the training of the student model by calculating the difference between the outputs of the teacher model and the student model; the interactive learning module enables the teacher model and the student model to transfer information to each other through joint training to achieve collaborative learning between the two; the classification and evaluation module makes a classification decision on the output features after feature extraction and knowledge distillation, and feeds back to other modules through accuracy evaluation to further optimize the system performance.

10. The knowledge distillation system based on the interaction and combination of ANN and SNN according to claim 9, wherein, The difference includes using KL divergence to measure the difference in the output distributions of the teacher model and the student model.

Citation Information

Cited By

  • Multi-scene image style migration and edge calculation optimization method

    CN120953099A