Object Tracking Method and Electronic Device Based on Pulse-Coded Learnable SNN
By transforming the SiamFC++ network to pulse encoding, the target tracking of video images is realized, the problem of single image recognition in the prior art cannot be tracked, the detection accuracy is improved, and it is suitable for real-time applications.
Patent Information
- Application Number
- CN202211084460.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-06
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2042-09-06
AI Technical Summary
In the prior art, the method based on pulsed neural network is only applicable to target recognition of a single image, and cannot be applied to target tracking of video images, and there is a lack of an effective target tracking solution.
The video to be tracked is decomposed into multiple frames of images, and the first frame and other frames are transformed at different scales. The modified SiamFC++ network is used for target tracking. The front-end network of this network includes two branches, each branch includes a pulse encoding module, a pulse feature extraction module and an ignition rate module. The image is encoded into a five-dimensional encoding vector through pulse encoding, and participates in the learning of pulse encoding parameters during the training process.
Real-time target tracking based on pulse coding is realized, the detection accuracy and computing efficiency are improved, and it is suitable for video surveillance and unmanned driving, and has application potential in artificial intelligence chips and neuromorphic hardware.
Smart Images

Figure CN115409870B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of spiking neural networks, and particularly relates to a target tracking method and an electronic device based on pulse-coded learnable SNN (Spiking Neural Network). Background Art
[0002] SNN is a new generation of artificial neural network model inspired by biology and belongs to a subset of deep learning. Different from the currently popular neural network and machine learning methods, SNN uses discrete pulses occurring at time points to transmit messages instead of common continuous values. SNN uses a model that fits the mechanism of biological neurons for calculation, which is closer to the biological neuron mechanism. Therefore, it has powerful capabilities such as spatio-temporal information representation, asynchronous event information processing, and network self-organization learning, and is expected to bridge the gap between neuroscience and machine learning, thereby further improving the operation effect of various tasks implemented based on neural networks.
[0003] In the prior art, the invention patent application with the publication number CN113111758A discloses a method for identifying ship targets in SAR (Synthetic Aperture Radar) images based on a spiking neural network; in this method, a visual saliency map extraction method based on a visual attention mechanism is adopted to obtain a visual saliency map from the SAR image; the visual saliency map is pulse-coded using a Poisson encoder to obtain a discrete pulse time series; a spiking neural network model is constructed using a convolutional neural network and LIF (Leaky Integrate-and-Fire) spiking neurons, and then the spiking neural network model is trained based on the obtained pulse time series, so as to accurately identify ship targets in the SAR image using the trained spiking neural network model.
[0004] However, the above-mentioned existing solution is only applicable to the target recognition of a single image and cannot be applied to the target tracking of video images; and in the prior art, there is no better solution for how to use SNN to achieve target tracking. Summary of the Invention
[0005] In order to solve the above problems existing in the prior art, the present invention provides a target tracking method and an electronic device based on pulse-coded learnable SNN.
[0006] The technical problems to be solved by the present invention are realized through the following technical solutions:
[0007] A target tracking method based on pulse-coded learnable SNN includes:
[0008] Decompose the video to be target-tracked into multiple frames of images, transform the first frame of the images to the first scale, and transform the other frames of the images to the second scale to obtain a preprocessed video frame sequence;
[0009] Perform target tracking on the video frame sequence based on a pre-trained spiking neural network model;
[0010] Among them, the spiking neural network model is obtained by transforming the SiamFC++ network. The transformation method is to replace the backbone of the SiamFC++ network with a front-end network; the front-end network includes two branches, and the two branches are respectively used for front-end processing of images of different scales;
[0011] Each branch of the front-end network includes a pulse encoding module, a pulse feature extraction module, and a firing rate calculation module. Among them, the pulse encoding module is used to encode the image into a five-dimensional encoding vector [t, b, c, h, w] composed of binary numbers; t represents the time dimension, b represents the batch size dimension, c represents the image channel dimension, h represents the image height dimension, and w represents the image width dimension; the pulse feature extraction module is used to extract pulse features from the five-dimensional encoding vector; the firing rate calculation module is used to remove the time dimension t from the pulse features.
[0012] Optionally, the pulse encoding module includes: a convolutional layer, a time domain expansion sub-module, and an IF neuron layer;
[0013] The convolutional layer is used to encode the image into a four-dimensional encoding vector [b, c, h, w]; the four-dimensional encoding vector is not composed of binary numbers;
[0014] The time domain expansion sub-module is used to expand the four-dimensional encoding vector in the time domain to obtain an expanded vector;
[0015] The IF neuron layer is used to activate the expanded vector into the five-dimensional encoding vector.
[0016] Optionally, the time domain expansion sub-module expands the four-dimensional encoding vector in the time domain to obtain an expanded vector, including: repeating the four-dimensional encoding vector T times to obtain an expanded vector; T ∈ t.
[0017] Optionally, an IF neuron is used as the activation function in the pulse feature extraction module.
[0018] Optionally, the training process of the spiking neural network model is as follows:
[0019] Obtain a video dataset; the video dataset includes multiple video samples, each video sample consists of multiple frames of images, and each frame of image is marked with a target position box;
[0020] Select a pair of images belonging to the same video sample from the video dataset as the template frame and the search frame respectively;
[0021] Input the template frame and the search frame into the two branches respectively, so that the spiking neural network model outputs a regression matrix, a classification matrix, and a quality matrix; wherein, the elements in the regression matrix are vectors representing the target tracking prediction box; the elements in the classification matrix are the classification results of whether the element corresponds to the target; the elements in the quality matrix are used to represent the prediction accuracy of the target tracking prediction box; the matrix dimensions of the regression matrix, the classification matrix, and the quality matrix are equal;
[0022] Evaluate whether the current spiking neural network model converges according to the regression matrix, the classification matrix, and the quality score matrix;
[0023] If not converged, adjust the network weight parameters of the current spiking neural network model, and select the next pair of images from the video dataset to continue training; otherwise, end the training and obtain the trained spiking neural network model.
[0024] Optionally, adjusting the network weight parameters of the current spiking neural network model includes: using the gradient descent method to adjust the network weight parameters of the current spiking neural network model;
[0025] The first gradient calculation formula is used in the gradient descent method to implement gradient calculation;
[0026] The first gradient calculation formula is:
[0027] Where x is the input of the activation function in the spiking neural network model, e is the natural logarithm base, and g'(x) represents the gradient.
[0028] Optionally, adjusting the network weight parameters of the current spiking neural network model includes: using the gradient descent method to adjust the network weight parameters of the current spiking neural network model;
[0029] The second gradient calculation formula is used in the gradient descent method to implement gradient calculation;
[0030] The second gradient calculation formula is:
[0031] x is the input of the activation function in the spiking neural network model, e is the natural logarithm base, and g'(x) represents the gradient.
[0032] Optionally, performing object tracking on the video frame sequence based on a pre-trained spiking neural network model, including:
[0033] Taking the first frame image in the video frame sequence as a template frame, and taking the i-th (i = [2, 3, 4, …, N]) frame image in the video frame sequence as a search frame, and inputting them into the trained spiking neural network model to obtain the classification matrix, regression matrix, and quality matrix output for the i-th time; N is the number of frames in the video frame sequence;
[0034] Multiplying the quality matrix and the classification matrix output for the i-th time to obtain the i-th classification score matrix;
[0035] According to the i-th classification score matrix, selecting the element with the highest corresponding classification score from the regression matrix output for the i-th time, and visually displaying the object tracking prediction box represented by this element in the i-th frame image.
[0036] Optionally, the video dataset includes: GOT10K.
[0037] In a second aspect, the present invention provides an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory complete communication with each other through the communication bus;
[0038] The memory is used to store a computer program;
[0039] The processor is used to implement the method steps of any of the above object tracking methods based on a spiking neural network model when executing the program stored in the memory.
[0040] In the object tracking method based on pulse-coded learnable SNN provided by the present invention, the SiamFC++ network capable of realizing object tracking is transformed. SiamFC++ is a CNN (Convolutional Neural Network). By replacing the network backbone at its front end with a front-end network, a pulsed neural network model capable of real-time object tracking is obtained. Among them, the front-end network includes two branches, and these two branches are respectively used for front-end processing of images of different scales. Each branch includes a pulse coding module, a pulse feature extraction module, and a firing rate calculation module. The pulse coding module is used to encode the image into a five-dimensional coding vector composed of binary numbers. The pulse feature extraction module is used to extract pulse features from the five-dimensional coding vector. The firing rate calculation module is used to eliminate the influence brought by the time dimension t from the pulse features. Compared with the prior art in which pulse coding is first performed outside the pulsed neural network and then the coding is input into the pulsed neural network for training, in the present invention, the pulse coding part participates in the training process of the pulsed neural network, thereby realizing the learnability of the parameters of the pulse coding part. Therefore, when object tracking is realized based on the trained pulsed neural network model, higher detection accuracy can be obtained.
[0041] Moreover, when the present invention performs pulse coding on the image, the time dimension is introduced, and a higher-dimensional five-dimensional coding vector is used to represent the image. Therefore, the pulse features extracted therefrom by the pulse feature extraction module can better express the effective information contained in the image. Although the firing rate calculation module is still used to remove the time dimension from the pulse features, the inventor found through comparative experiments that the method of introducing the time dimension and then removing it has better performance than the method of directly not introducing the time dimension.
[0042] The following will further elaborate on the present invention in conjunction with the accompanying drawings. Description of the Drawings
[0043] Figure 1 is a flowchart of an object tracking method based on pulse-coded learnable SNN provided by an embodiment of the present invention;
[0044] Figure 2 is a schematic structural diagram of a pulsed neural network model used in an embodiment of the present invention;
[0045] Figure 3 is Figure 2 a schematic structural diagram of the front-end network in the shown pulsed neural network model;
[0046] Figure 4 is a schematic diagram of realizing pulse coding in an embodiment of the present invention;
[0047] Figure 5It is the visualization effect diagram after visualizing the quality score matrix and the classification score matrix in the embodiments of the present invention;
[0048] Figure 6 It is the effect diagram of target tracking using a target tracking method based on a pulse-coded learnable SNN provided by the embodiments of the present invention;
[0049] Figure 7 It is the structural schematic diagram of an electronic device provided by the embodiments of the present invention. Specific embodiments
[0050] The following further describes the present invention in detail with specific embodiments, but the embodiments of the present invention are not limited thereto.
[0051] In order to use an SNN to achieve target tracking, the embodiments of the present invention provide a target tracking method based on a pulse-coded learnable SNN,
[0052] As Figure 1 shown, the method includes the following steps:
[0053] S10: Decompose the video to be target-tracked into multiple frames of images, transform the first frame of the images to a first scale, and transform the other frames of the images to a second scale to obtain a preprocessed video frame sequence.
[0054] Specifically, the first frame of the images can be transformed to the first scale and the other frames of the images can be transformed to the second scale by using the method of central cropping or random cropping, and the cropping is performed with reference to the position of the target as the center during the cropping process.
[0055] Exemplarily, the first scale can be 127×127, and the second scale can be 303×303, but it is not limited thereto.
[0056] S20: Use the pre-trained spiking neural network model to perform target tracking on the video frame sequence.
[0057] In this step, the spiking neural network model is a spiking neural network model obtained by modifying the SiamFC++ network. As Figure 2 shown, the modification method is to replace the network backbone in the SiamFC++ network with a front-end network; the front-end network includes two branches, and these two branches are respectively used for front-end processing of images with different scales.
[0058] In the prior art, in addition to SiamFC++, there are also SiamFC and SiamRPN++ which are tracking algorithms based on Siamese networks. Among them, SiamFC can maintain high tracking accuracy while having a relatively fast tracking speed; SiamRPN++ successfully breaks the spatial invariance limit, trains a Siamese tracker based on the Resnet architecture, introduces a deep network into object tracking, and achieves good results; in comparison, SiamFC++ is a network architecture based on an anchor-free model, and its tracking performance is more excellent, with the FPS (frames per second) reaching 160 frames per second.
[0059] Among them, combined with Figure 2 and 3 It can be seen that each branch of the front-end network includes a pulse coding module, a pulse feature extraction module, and a firing rate calculation module. Among them, the pulse coding module is used to encode an image into a five-dimensional coding vector [t, b, c, h, w] composed of binary numbers; t represents the time dimension, b represents the batchsize dimension, c represents the image channel dimension, h represents the image height dimension, and w represents the image width dimension; the pulse feature extraction module is used to extract pulse features from the five-dimensional coding vector; the firing rate calculation module is used to eliminate the influence brought by the time dimension t from the pulse features.
[0060] Specifically, the above-mentioned pulse coding module may include: a convolutional layer, a time domain expansion sub-module, and an IF neuron layer.
[0061] Among them, the convolutional layer is used to encode an image into a four-dimensional coding vector [b, c, h, w]; this four-dimensional coding vector is not composed of binary numbers.
[0062] The time domain expansion sub-module is used to expand the four-dimensional coding vector in the time domain to obtain an expanded vector;
[0063] Exemplarily, the time domain expansion sub-module can repeat the four-dimensional coding vector T times to obtain an expanded vector; T ∈ t. Preferably, T can be taken as 6, but it is not limited thereto.
[0064] In another implementation, the convolutional layer can be made to repeatedly encode the image T times. Correspondingly, the time domain expansion sub-module can splice the T four-dimensional coding vectors obtained from these T encodings to obtain an expanded vector.
[0065] It should be noted that the time domain expansion sub-module does not increase the number of weight parameters of the spiking neural network.
[0066] The IF neuron layer is used to activate the expanded vector into a five-dimensional coding vector composed of binary numbers.
[0067] Figure 4 An exemplary schematic diagram of the pulse coding process is shown, where each pixel point of the input image is encoded into a pulse sequence of 96 with a length of T.
[0068] In another implementation, the IF neuron layer in the pulse coding module can be replaced with an LIF neuron layer.
[0069] The pulse feature extraction module can adopt the image feature extraction network in the existing CNN network; or, in order to make the pulse neural network closer to biological characteristics, the activation function in these image feature extraction networks can be replaced with an IF neuron or an LIF neuron. Thus, when the pulse neural network performs forward propagation, neurons are used instead of activation functions, and the accumulation of membrane voltage and the firing of pulses are completed through neurons, thereby realizing the non-linear calculation of the network.
[0070] Exemplarily, Table 1 gives the parameters of some network structures in the front-end network.
[0071] Table 1
[0072]
[0073] Among them, Conv2d all represents the convolutional layer, IFNode all represents the IF neuron layer, and MaxPool2d all represents the max pooling layer. Module represents the module, Layer represents the level in the module, Param represents the number of weights involved in the level, z’shape represents the feature dimension output by each level in the branch where the template frame enters. For example, the feature dimension output by the Conv2d in the first row of Table 1 is T×B×96×59×59; where T is the specific size of the time dimension, and B represents the size of the batch size. x’shape represents the feature dimension output by each level in the branch where the search frame enters. Among them, the template frame refers to the video frame used as a template to search for the target to achieve target tracking during target tracking, and the search frame is the video frame from which the target needs to be found.
[0074] After the front-end network, it is the remaining back-end network of the SiamFC++ network after removing its backbone; this back-end network includes a convolutional layer dedicated to the classification sub-task and a convolutional layer dedicated to the regression sub-task Both groups of pulse features output by the front-end network are sent into and and Both output two feature maps; here, The two feature maps output by and are represented by The two feature maps output by and denotes; where ζ represents the pulse coding module, φ represents the pulse feature extraction module, represents the convolutional layer dedicated to the subtasks of the backend network. Z represents the template frame, and X represents the search frame.
[0075] Then, perform a cross-correlation operation on the two output feature maps, that is the result is denoted by ξ cls ; and, perform a cross-correlation operation on the two output feature maps, that is the result is denoted by ξ reg The results of the two cross-correlation operations respectively pass through 3 convolutional layers with a kernel size of 3×3; then, the output of ξ cls after passing through the 3×3 convolutional layer is respectively fed into two convolutional layers with a kernel size of 1×1, and the classification matrix and the quality matrix are respectively output by these two convolutional layers; in addition, the output of ξ reg after passing through the 3×3 convolutional layer is fed into a convolutional layer with a kernel size of 1×1, and the regression matrix is output by this convolutional layer.
[0076] Among them, the elements in the regression matrix are used to represent the target tracking prediction box; specifically, the elements stored in the regression matrix are the distance values corresponding to the up, down, left, and right four directions respectively, and the position of this element has a point mapping relationship with the elements in the search frame. After mapping this element back to the search frame, taking the element mapped in the search frame as the center, expand the above four distance values in the up, down, left, and right four directions respectively, and a target tracking prediction box can be obtained; the elements in the classification matrix are the classification results of whether this element corresponds to the target; the elements in the quality matrix are used to represent the prediction accuracy of the target tracking prediction box; the matrix dimensions of the regression matrix, the classification matrix, and the quality matrix are equal.
[0077] After building the spiking neural network model, it can be trained, and there are multiple training methods. Exemplarily, in one implementation, the training process of the spiking neural network model can be as follows:
[0078] (1) Obtain a video dataset; the video dataset includes multiple video samples, each video sample consists of multiple frames of images, and each frame of image is marked with a target position box.
[0079] (2) Select a pair of images belonging to the same video sample from the video dataset as the template frame and the search frame respectively.
[0080] (3) Input the template frame and the search frame into the two branches of the front-end network respectively, so that the spiking neural network model outputs the regression matrix, the classification matrix, and the quality matrix.
[0081] (4) Evaluate whether the current spiking neural network model converges according to the regression matrix, classification matrix, and quality matrix;
[0082] Specifically, calculate the loss value of the model according to the regression matrix, classification matrix, and quality matrix; if the loss value converges, the spiking neural network model converges, otherwise the spiking neural network model does not converge.
[0083] Among them, the calculation formula of the loss value is:
[0084]
[0085] Among them, p x,y represents the element at the point (x, y) in the classification matrix, represents the marking value indicating whether the pixel point having a mapping relationship with this point (x, y) in the search frame is located in the target position box of the search frame, where represents being located, represents not being located, N pos represents the number of. Among them, the position where the pixel points having a mapping relationship with the points (x, y) of the classification matrix or regression matrix in the search frame is S is the total stride of the front-end network.
[0086] q x,y represents the element at the point (x, y) in the quality matrix; t x,y represents the element at the point (x, y) in the regression matrix, represents the intersection over union of the target tracking prediction box characterized by t x,y and the target position box of the search frame; the calculation formula of this intersection over union is expressed as:
[0087]
[0088] Among them, B represents the target tracking prediction box, B * represents the target position box, Intersection(·) represents finding the intersection, and Union(·) represents finding the union.
[0089] is the vector composed of the distances between the target tracking prediction box characterized by t x,y and the target position box of the search frame in the up, down, left, and right four directions, and the calculation method of is:
[0090]
[0091]
[0092]
[0093]
[0094] Among them, (x0, y0) and (x1, y1) are the coordinates of the upper left corner and the lower right corner of the target position box respectively. By l * ,t * 、r * , b * The vector composed of these four distances (l * ,t * ,r * ,b * ).
[0095] L represents the loss value; L cls (·) is the Focal Loss function, L quality (·) is the BCE loss function, L reg is the IoU loss function.
[0096] In the embodiment of the present invention, since the intersection-over-union ratio of the target position frame and the target tracking prediction frame is introduced into the loss function used to calculate the loss value, the accuracy of the trained spiking neural network model in target tracking can be improved.
[0097] (5) If the model has not converged, the network weight parameters of the current spiking neural network model are adjusted, and the next pair of images are selected from the video data set to continue training; otherwise, the training is terminated to obtain a trained spiking neural network model.
[0098] Compared with mature artificial neural network (ANN) training algorithms, one of the most difficult problems in spiking neural network research is the difficulty in training due to complex dynamics and the non-differentiable characteristics of pulses. Specifically, during the forward propagation of the training process, the spiking neuron uses a step function, and the output value is discrete 0 or 1. However, in the back propagation process, due to the non-differentiable characteristics of the spiking neuron, it is impossible to use back propagation to train the spiking neural network normally. To this problem, researchers have tried to fundamentally alleviate it and want to optimize network performance by strengthening SNN. The main optimization directions include two categories: one is to achieve optimization by satisfying known biological discoveries as much as possible with the ultimate goal of understanding biological systems, such as synaptic plasticity dependent on pulse timing, long-term potentiation, long-term inhibition, neuronal lateral inhibition, etc.; the other is to target computational performance, such as gradient substitution, ANN to SNN conversion, pulse smoothing approximation, etc.
[0099] In the embodiment of the present invention, during backpropagation, the gradient descent method is used to adjust the network weight parameters of the current spiking neural network model. The optimizer adjusts the weights according to the gradient of the loss function at the current position, thereby completing the optimization of the model. Moreover, in the embodiment of the present invention, a special substitution function is used to solve the problem of difficult direct training caused by the non-differentiable characteristics of spiking neurons.
[0100] Among them, in the gradient descent method, the first gradient calculation formula can be used to implement gradient calculation; the first gradient calculation formula is:
[0101]
[0102] Among them, x represents the input of the activation function in the spiking neural network model, e is the natural base, and g'(x) represents the gradient.
[0103] Alternatively, in the gradient descent method, the second gradient calculation formula can also be used to implement gradient calculation; the second gradient calculation formula is:
[0104]
[0105] Among them, x also represents the input of the activation function in the spiking neural network model, e is the natural base, and g'(x) represents the gradient.
[0106] After the training is completed, save the weights obtained from the current training, and then the target tracking can be performed based on the trained spiking neural network model.
[0107] Specifically, performing target tracking on the video frame sequence based on the trained spiking neural network model includes:
[0108] (1) Take the first frame image in the video frame sequence obtained in step S10 as the template frame, and take the i-th (i = [2, 3, 4, L, N]) frame image in the video frame sequence as the search frame and input it into the trained spiking neural network model to obtain the classification matrix, regression matrix, and quality matrix output for the i-th time; N is the number of frames in the video frame sequence.
[0109] It can be understood that each time only one frame in i = [2, 3, 4, L, N] is used as the search frame and input into one branch of the front-end network, and at the same time, the first frame is used as the template frame and input into another branch of the front-end network.
[0110] (2) Multiply the quality matrix and the classification matrix output for the i-th time to obtain the i-th classification score matrix.
[0111] Figure 5 Subfigure (a) in shows a classification matrix with an image output. Figure 5Sub - figure (b) shows a quality score matrix with the output being an image. The two are multiplied to obtain a classification score matrix. It can be understood that multiplying the quality score matrix by the classification matrix, the resulting classification score matrix can help find the target more precisely.
[0112] (3) According to the i - th classification score matrix, select the element with the highest corresponding classification score from the i - th output regression matrix, and visually display the target tracking prediction box represented by this element in the i - th frame image.
[0113] Specifically, assume that the value of the point (x, y) in the i - th new classification score matrix is the largest, that is, the classification score of this point is the highest. Then select the point (x, y) from the i - th output regression matrix as well, and visually display the target tracking prediction box represented by this point in the i - th frame image.
[0114] Next, simulation experiments are used to further illustrate the beneficial effects of the embodiments of the present invention.
[0115] First, build the front - end network according to Table 1, and correspondingly build the spiking neural network as Figure 2 shown. In this spiking neural network, the template frame size is set to 127×127, and the search frame size is set to 303×303. The spike encoding module of the spiking neural network performs spike encoding on the template frame and the search frame; the encoded five - dimensional encoding vectors are respectively input into the spike feature extraction module for feature extraction. Set the observation time as T. Then, for the template frame, a spike feature of T×256×6×6 is extracted, and for the input of the search frame, a spike feature of T×256×28×28 is extracted; then the firing rate operation is performed to obtain spike features of 256×6×6 and 256×28×28 respectively. For the two spike features obtained after the firing rate operation, they enter the back - end network for processing to obtain a 1×17×17 classification matrix and a 4×17×17 regression matrix;
[0116] Then the parameters are initialized, including: setting the video dataset retrieval path, equipment, observation time, training cycle, output path, etc. Among them, the video dataset uses GOT10K, which is released by the Chinese Academy of Sciences and includes more than 10,000 video clips taken from the real world and more than 560 categories, and more than 1.5 million manually marked target position boxes. The equipment used includes two CPUs (Central Processing Units, central processing units), model Intel (R) Xeon (R) CPUE5-2620v4, running frequency of 2.10GHz, memory capacity of 64G, frequency of 2400MHz, and two GPUs (Graphics Processing Units, graphics processing units), model Nvidia RTX 2070Super, video memory of 8G×2. Observation time T = 6; training cycle Epochs = 20. The output path is the path for saving the test results, which can be selected according to the actual situation. In addition, the software environment of this simulation experiment is implemented in Python language, and the deep learning framework is PyTorch, version 1.8.0.
[0117] Then, the pulse neural network is trained, wherein each time an image is input, the loss value is calculated based on the classification matrix and regression matrix output by the model, and the network weight parameters are adjusted and optimized accordingly until the loss value converges to obtain a trained pulse neural network model.
[0118] After training the spiking neural network model, the model is tested using the Basketball sequence in the OTB100 dataset. The specific process is as follows:
[0119] The first frame in the Basketball sequence is used as the template frame, and the other frames are used as search frames in turn. The template frame and search frame are respectively input into the trained spiking neural network model so that the model outputs the classification matrix, regression matrix and quality matrix. The quality matrix is multiplied by the classification matrix to obtain the classification score matrix, and the obtained classification score matrix is used to select the target position prediction box to achieve target tracking. The target tracking effect is shown in Figure 1. Figure 5 As shown in the figure, the dark box is the visualization result of the selected target position prediction box, and the light box is the visualization result of the target position box annotated in the image itself in the Basketball sequence. In addition, Figure 5 The upper left corner shows the sequence number of the image in the video sequence, and the lower right corner shows the intersection-over-union ratio between the target position prediction box and the target position box in the current frame.
[0120] In the object tracking method based on pulse-coded learnable SNN provided by the embodiments of the present invention, the SiamFC++ network capable of achieving object tracking is transformed. SiamFC++ is a CNN (Convolutional Neural Network). By replacing the network backbone at its front end with a front-end network, a pulse neural network model capable of real-time object tracking is obtained. Among them, the front-end network includes two branches, which are respectively used for front-end processing of images of different scales. Each branch includes a pulse coding module, a pulse feature extraction module, and a firing rate calculation module. The pulse coding module is used to encode the image into a five-dimensional coding vector composed of binary numbers. The pulse feature extraction module is used to extract pulse features from the five-dimensional coding vector. The firing rate calculation module is used to eliminate the influence brought by the time dimension t from the pulse features. Compared with the prior art in which pulse coding is first performed outside the pulse neural network and then the encoded data is input into the pulse neural network for training, in the embodiments of the present invention, the pulse coding part participates in the training process of the pulse neural network, so that the parameters of the pulse coding part can be learned. Therefore, when object tracking is realized based on the trained pulse neural network model, higher recognition accuracy can be obtained.
[0121] Moreover, in the embodiments of the present invention, the time dimension is introduced when the image is pulse-coded, and a higher-dimensional five-dimensional coding vector is used to represent the image. Therefore, the pulse features extracted by the pulse feature extraction module from it can better express the effective information contained in the image. Although the firing rate calculation module is still used to remove the time dimension from the pulse features, the inventor found through comparative experiments that the method of introducing the time dimension and then removing it has better performance than the method of directly not introducing the time dimension.
[0122] In addition, the embodiments of the present invention also have the following beneficial effects:
[0123] The embodiments of the present invention use discrete binary sequences for information transmission and calculation, and have a faster inference speed than traditional artificial neural networks during the inference process, high computational efficiency, can perform real-time object tracking, and can be used in specific fields such as video surveillance and unmanned driving in practice.
[0124] The embodiments of the present invention provide a feasible solution for transplanting SNN to edge devices such as artificial intelligence chips or neuromorphic hardware.
[0125] Based on the same inventive concept, the embodiments of the present invention also provide an electronic device, as Figure 7 shown, including a processor 701, a communication interface 702, a memory 703, and a communication bus 704. Among them, the processor 701, the communication interface 702, and the memory 703 complete mutual communication through the communication bus 704.
[0126] A memory 703 for storing a computer program;
[0127] A processor 701, when executing the program stored on the memory 703, implements the method steps described in any of the above target tracking methods based on pulse-coded learnable SNNs.
[0128] The communication bus mentioned in the above electronic device can be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, only a thick line is used in the figure to represent it, but it does not mean that there is only one bus or one type of bus.
[0129] The communication interface is used for communication between the above electronic device and other devices.
[0130] The memory may include a Random Access Memory (RAM), and may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0131] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processing (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0132] In practical applications, the above electronic device may be a desktop computer, a portable computer, a smart mobile terminal, a server, etc. There is no limitation here, and any electronic device that can implement the present invention belongs to the protection scope of the present invention.
[0133] The present invention also provides a computer-readable storage medium. A computer program is stored in the computer-readable storage medium. When the computer program is executed by a processor, the method steps described in any of the above-mentioned object tracking methods based on pulse-coded learnable SNN are implemented.
[0134] Optionally, the computer-readable storage medium may be a non-volatile memory (Non-Volatile Memory, NVM), such as at least one disk memory.
[0135] Optionally, the computer-readable memory may also be at least one storage device located away from the aforementioned processor.
[0136] In another embodiment of the present invention, a computer program product containing instructions is also provided. When it runs on a computer, it causes the computer to execute the method steps described in any of the above-mentioned object tracking methods based on pulse-coded learnable SNN.
[0137] It should be noted that for the embodiments of the electronic device / storage medium / computer program product, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiments.
[0138] It should be noted that the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the present invention, "a plurality" means two or more, unless otherwise specifically defined.
[0139] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, those skilled in the art can combine and combine the different embodiments or examples described in this specification.
[0140] Although the present invention has been described in conjunction with various embodiments herein, however, in the process of implementing the claimed present invention, those skilled in the art can understand and achieve other variations of the disclosed embodiments by viewing the accompanying drawings, the disclosure content, and the appended claims.
[0141] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (devices), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or a device for implementing the functions specified in multiple blocks.
[0142] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or the functions specified in multiple blocks.
[0143] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or the functions specified in multiple blocks.
[0144] The above content is a further detailed description of the present invention in combination with specific preferred embodiments. It cannot be determined that the specific implementation of the present invention is only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention pertains, without departing from the concept of the present invention, several simple deductions or substitutions can be made, and all should be regarded as belonging to the protection scope of the present invention.
Claims
1. A target tracking method based on a pulse-coded learnable SNN, characterized in that, Including: Decompose the video to be target-tracked into multiple frames of images, transform the first frame of the images to the first scale, and transform the other frames of the images to the second scale, to obtain a preprocessed video frame sequence; Perform target tracking on the video frame sequence based on a pre-trained spiking neural network model; Wherein, the spiking neural network model is obtained by modifying the SiamFC++ network, and the modification method is to replace the backbone of the SiamFC++ network with a front-end network; the front-end network includes two branches, and the two branches are respectively used for front-end processing of images of different scales; Each branch of the front-end network includes a pulse encoding module, a pulse feature extraction module, and a firing rate calculation module. Among them, the pulse encoding module is used to encode the image into a five-dimensional encoding vector [t, b, c, h, w] composed of binary numbers; t represents the time dimension, b represents the batch size dimension, c represents the image channel dimension, h represents the image height dimension, and w represents the image width dimension; the pulse feature extraction module is used to extract pulse features from the five-dimensional encoding vector; the firing rate calculation module is used to remove the time dimension t from the pulse features.
2. The target tracking method based on a pulse-coded learnable SNN according to claim 1, wherein, The pulse encoding module includes: a convolutional layer, a time domain expansion sub-module, and an IF neuron layer; The convolutional layer is used to encode the image into a four-dimensional encoding vector [b, c, h, w]; the four-dimensional encoding vector is not composed of binary numbers; The time domain expansion sub-module is used to expand the four-dimensional encoding vector in the time domain to obtain an expanded vector; The IF neuron layer is used to activate the expanded vector into the five-dimensional encoding vector.
3. The target tracking method based on pulse-coded learnable SNN according to claim 2, wherein The time domain expansion sub-module expands the four-dimensional encoding vector in the time domain to obtain an expanded vector, including: repeating the four-dimensional encoding vector T times to obtain an expanded vector; T ∈ t.
4. The target tracking method based on a pulse-coded learnable SNN according to claim 1, wherein The IF neuron is used as the activation function in the pulse feature extraction module.
5. The target tracking method based on pulse-coded learnable SNN according to claim 1, wherein, The training process of the spiking neural network model is as follows: Obtain a video data set; the video data set includes multiple video samples, each video sample is composed of multiple frames of images, and each frame of image is marked with a target position box; Select a pair of images belonging to the same video sample from the video data set as the template frame and the search frame respectively; Input the template frame and the search frame into the two branches respectively, so that the spiking neural network model outputs a regression matrix, a classification matrix, and a quality matrix; wherein, the elements in the regression matrix are vectors representing the target tracking prediction box; the elements in the classification matrix are the classification results of whether the element corresponds to the target; the elements in the quality matrix are used to represent the prediction accuracy of the target tracking prediction box; the matrix dimensions of the regression matrix, the classification matrix, and the quality matrix are equal; Evaluate whether the current spiking neural network model converges according to the regression matrix, the classification matrix, and the quality matrix; If not converged, adjust the network weight parameters of the current spiking neural network model, and select the next pair of images from the video dataset for continued training; otherwise, end the training and obtain the trained spiking neural network model.
6. The target tracking method based on pulse-coded learnable SNN according to claim 5, wherein, The adjustment of the network weight parameters of the current spiking neural network model includes: using the gradient descent method to adjust the network weight parameters of the current spiking neural network model; In the gradient descent method, the first gradient calculation formula is used to implement gradient calculation; The first gradient calculation formula is as follows: Among them, x is the input of the activation function in the spiking neural network model, e is the natural base, and g'(x) represents the gradient.
7. The target tracking method based on a pulse-coded learnable SNN according to claim 5, wherein The adjustment of the network weight parameters of the current spiking neural network model includes: using the gradient descent method to adjust the network weight parameters of the current spiking neural network model; In the gradient descent method, the second gradient calculation formula is used to implement gradient calculation; The second gradient calculation formula is as follows: x is the input of the activation function in the spiking neural network model, e is the natural base, and g'(x) represents the gradient.
8. The target tracking method based on pulse-coded learnable SNN according to claim 5, characterized in that Target tracking of the video frame sequence based on the pre-trained spiking neural network model includes: Taking the first frame image in the video frame sequence as the template frame and taking the i-th (i = [2, 3, 4, L, N]) frame image in the video frame sequence as the search frame and inputting them into the trained spiking neural network model to obtain the classification matrix, regression matrix, and quality matrix output for the i-th time; N is the number of frames in the video frame sequence; Multiply the quality matrix and the classification matrix output for the i-th time to obtain the i-th classification score matrix; According to the i-th classification score matrix, select the element with the highest corresponding classification score from the regression matrix output for the i-th time, and visually display the target tracking prediction box represented by this element in the i-th frame image.
9. The target tracking method based on pulse-coded learnable SNN according to claim 5, wherein, The video dataset includes: GOT10K.
10. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used to store computer programs; When the processor is used to execute the computer program stored on the memory, it implements the method steps of the target tracking method based on the pulse-coded learnable SNN described in any one of claims 1 to 9.
Citation Information
Patent Citations
SAR image ship target identification method based on pulse neural network
CN113111758A
Target tracking method and system based on pulse convolutional neural network
CN114694079A
Target tracking method, target tracking device and computer readable medium
CN114926505A