A target tracking method and system based on pulse convolutional neural network

Through the method based on pulse convolution neural network, video sequences based on image frames and event frames are directly processed, and the problem of difficult to process event frames and poor tracking in complex environments in the prior art is solved, thereby achieving efficient and accurate target tracking.

CN114694079BActive Publication Date: 2025-05-13ZHEJIANG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210407708.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-19
Publication Date
2025-05-13
Estimated Expiration
2042-04-19

AI Technical Summary

Technical Problem

Existing target tracking algorithms are difficult to process event frame-based video sequences without preprocessing, and tracking is poor in environments with complex backgrounds and noise events.

Method used

The target tracking method based on pulse convolution neural network is used to process the target video sequence through the trained pulse convolution neural network model, including the convolution layer, the flat layer and the fully connected layer, and the neuron model of the current-based leakage integration distribution model is used. This method can directly process video sequences based on image frames and event frames, extract complex features and achieve accurate target tracking.

Benefits of technology

It realizes efficient and accurate target tracking without preprocessing, can handle complex backgrounds and many noise events, and improves the stability and accuracy of tracking.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114694079B_ABST
    Figure CN114694079B_ABST
Patent Text Reader

Abstract

The present invention discloses a target tracking method based on a pulse convolutional neural network, and relates to the technical field of target tracking. The method comprises: inputting a target video sequence to be identified into a pulse convolutional neural network model, determining the classification result of the target area and the classification result of the background area; the classification result of the target area is used to determine the driving track of the target; the target video sequence is a video sequence based on an image frame or a video sequence based on an event frame; an image frame and an event frame both include a target area and a background area; the pulse convolutional neural network model comprises at least a trained first pulse convolutional neural network; the first pulse convolutional neural network comprises a first convolutional layer, a second convolutional layer, a third convolutional layer, a flat layer, a first fully connected layer and a second fully connected layer connected in sequence; the model of the neuron in the first pulse convolutional neural network is a leakage integrated release model based on current. The present invention can accurately and efficiently track targets.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of target tracking technology, and in particular to a target tracking method and system based on a pulse convolutional neural network. Background Art

[0002] With the rapid development of computer vision, traditional cameras are gradually unable to meet people's needs in certain specific environments (such as high-speed movement, too dark or too much environment). In recent years, researchers have developed a series of event cameras through the exploration of biological retinas. Compared with traditional cameras, event cameras do not output frame images at a specific frequency, but encode the dynamic visual information of the scene into a continuous spatiotemporal pulse event stream, which has the characteristics of high dynamic range, high temporal resolution, low power consumption and low information redundancy.

[0003] Target tracking is an important problem in the field of computer vision. Currently, most tracking algorithms are based on traditional image frames, which can be divided into two categories: tracking based on correlation filtering and tracking based on traditional deep neural networks. Bolme et al. first introduced correlation filtering into the problem of target tracking. They proposed a minimum output sum of squared error (MOSSE) filter, but the tracking performance was poor. To solve this problem, Henriques et al. proposed kernelized correlation filters (KCF), which use the gradient characteristics of multi-channel histograms. However, the features extracted by the filter are relatively superficial, which makes it easy to lose the target during the tracking process.

[0004] In order to extract more complex features, trackers in recent years have been developed based on traditional deep neural networks. Nam et al. proposed the MDNet network, which includes offline multi-domain learning and online specific domain learning. The Siamrpn++ network proposed by Li et al. deepens the depth of the twin network and greatly improves the tracking accuracy. The STARK network proposed by Yan et al. learns a robust spatiotemporal joint representation through Transformer, models target tracking as a direct bounding box prediction problem, and has achieved state-of-the-art performance on many data sets. However, these trackers based on traditional deep neural networks cannot directly process event-frame-based video sequences, and they need to be converted into frames first, and some temporal information will be lost in the process.

[0005] Currently, there are few trackers that directly process event input and indicate tracking results through bounding boxes. The ECT and RCT algorithms encapsulated in the jAER software are essentially clustering algorithms, which can achieve better tracking in scenes where only the target is moving. Once the background is too complex, the texture is rich, and there are too many noise events, the tracking results will deviate from the target. Summary of the invention

[0006] The purpose of the present invention is to provide a target tracking method and system based on a pulse convolutional neural network, which can achieve target tracking efficiently and accurately.

[0007] To achieve the above object, the present invention provides the following solutions:

[0008] A target tracking method based on a pulse convolutional neural network, comprising:

[0009] Acquire a target video sequence to be identified; the target video sequence is a video sequence based on an image frame or a video sequence based on an event frame; one of the image frames and one of the event frames both include a target area and a background area;

[0010] Inputting the target video sequence to be identified into the pulse convolutional neural network model to determine the classification result of the target area and the classification result of the background area; the classification result of the target area is used to determine the driving trajectory of the target;

[0011] The pulse convolutional neural network model includes at least one trained first pulse convolutional neural network; the first pulse convolutional neural network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a flat layer, a first fully connected layer and a second fully connected layer connected in sequence; the model of neurons in the first pulse convolutional neural network is a current-based leakage integration and release model.

[0012] Optionally, when the target video sequence includes a video sequence based on image frames and a video sequence based on event frames, the pulse convolutional neural network model further includes a trained second pulse convolutional neural network with the same structure as the first pulse convolutional neural network;

[0013] The first pulse convolutional neural network is used to determine the classification result of the target area and the classification result of the background area according to the video sequence based on the image frame;

[0014] The second pulse convolutional neural network is used to determine the classification result of the target area and the classification result of the background area according to the event frame-based video sequence.

[0015] Optionally, inputting the target video sequence to be identified into a pulse convolutional neural network model to determine the classification result of the target area and the classification result of the background area specifically includes:

[0016] Preprocessing the target video sequence to be identified;

[0017] Fine-tuning the parameters of the spike convolutional neural network model according to the target area and the background area of ​​the first frame of the preprocessed target video sequence; the first frame is an image frame or an event frame;

[0018] The preprocessed target video sequence is input into the fine-tuned pulse convolutional neural network model to determine the classification result of the target area and the classification result of the background area.

[0019] Optionally, the target video sequence to be identified is preprocessed, specifically including:

[0020] When the target video sequence is a video sequence based on image frames, normalizing the target video sequence to be identified to obtain a preprocessed target video sequence;

[0021] When the target video sequence is a video sequence based on event frames, the target video sequence to be identified is subjected to horizontal and vertical coordinate reduction processing to obtain a preprocessed target video sequence.

[0022] Optionally, fine-tuning the parameters of the spiking convolutional neural network model according to the target area and the background area of ​​the first frame of the preprocessed target video sequence specifically includes:

[0023] According to the first frame of the preprocessed target video sequence, a fine-tuning sample set is obtained; the fine-tuning sample set includes a plurality of positive sample frames randomly generated near the target bounding box of the first frame using a normal distribution algorithm, a plurality of first negative sample frames randomly generated near the target bounding box of the first frame using a uniform distribution algorithm, and a plurality of second negative sample frames randomly generated within the entire image range of the first frame using a uniform distribution algorithm; wherein an intersection-over-union ratio of the positive sample frame to the target bounding box is greater than a first threshold; and an intersection-over-union ratio of the negative sample frame to the target bounding box is less than a second threshold;

[0024] The parameters of the pulse convolutional neural network model are fine-tuned according to the fine-tuning sample set.

[0025] Optionally, the step of inputting the preprocessed target video sequence into the fine-tuned pulse convolutional neural network model to determine the classification result of the target area and the classification result of the background area specifically includes:

[0026] According to the first frame of the preprocessed target video sequence, a regressor sample set is obtained; the regressor sample set includes a plurality of training sample frames randomly generated around the target bounding box of the first frame according to a uniform distribution algorithm; the intersection and union ratio of the training sample frame and the target bounding box is greater than a third threshold;

[0027] Inputting the regressor sample set into the fine-tuned spiking convolutional neural network model so that the features of the third convolutional layer output in the fine-tuned spiking convolutional neural network model are used to train a bounding box regressor to obtain a trained bounding box regressor;

[0028] Inputting the preprocessed target video sequence into the fine-tuned pulse convolutional neural network model to obtain the classification result of the target area and the classification result of the background area; inputting the marking feature output by the third convolutional layer in the fine-tuned pulse convolutional neural network model into the trained border regressor to obtain the correction result of the target area; the marking feature is the feature output by the third convolutional layer after the preprocessed target video sequence is input into the fine-tuned pulse convolutional neural network model;

[0029] The classification result of the target area is classified according to the correction result of the target area to obtain a corrected classification result of the target area.

[0030] Optionally, the first convolution layer includes 96 7×7 filters; the stride of the first convolution layer is 2; the input of the first convolution layer is a matrix of batch×C×107×107, and the output of the first convolution layer is a binary pulse matrix of batch×96×25×25; wherein batch represents the number of batches input into the first pulse convolutional neural network or the second pulse convolutional neural network at one time; when the target video sequence is a video sequence based on image frames, C=3; when the target video sequence is a video sequence based on event frames, C=1;

[0031] The second convolutional layer includes 256 5×5 filters; the stride of the second convolutional layer is 2; the input of the second convolutional layer is a binary pulse matrix of batch×96×25×25, and the output of the second convolutional layer is a binary pulse matrix of batch×256×11×11;

[0032] The third convolutional layer includes 512 3×3 filters; the step size of the third convolutional layer is 1; the input of the third convolutional layer is a binary pulse matrix of batch×96×25×25, and the output of the third convolutional layer is a binary pulse matrix of batch×512×3×3;

[0033] The flattening layer is used to stretch the input batch×512×3×3 binary pulse matrix into a batch×4608 binary pulse vector;

[0034] The input of the first fully connected layer is a binary pulse vector of batch×4608, and the output of the first fully connected layer is a binary pulse vector of batch×512;

[0035] The input of the second fully connected layer is a binary pulse vector of batch×512, and the output of the second fully connected layer is a binary pulse vector of batch×2; the number of neurons output by the second fully connected layer is 2, wherein one neuron is used to output the classification result of the target area, and the other neuron is used to output the classification result of the background area.

[0036] Optionally, the kinetic equation of the current-based leakage integration release model is:

[0037] V t =L t -E t

[0038]

[0039]

[0040]

[0041] Among them, V t represents the membrane voltage of the neuron at time t, L t represents the membrane voltage calculated by the kinetic equation of the leak-integration-spiking model at time t, E t represents the reset item at time t; N represents N input neurons, w i represents the synaptic weight of the i-th input neuron, represents the moment when the i-th input neuron in the previous layer emits the j-th output pulse, represents the time when the sth input neuron in the current layer emits the jth output pulse; V0 represents the normalization factor; α m , α s Both represent learnable time decay factors, θ represents the threshold of the neuron; K represents the standardized postsynaptic potential kernel.

[0042] Optionally, the first trained pulse convolutional neural network and the second trained pulse convolutional neural network are both determined after training the pulse convolutional neural network using a training set and a spatiotemporal credit allocation learning algorithm;

[0043] The training set includes a plurality of sample pairs; each of the sample pairs includes a positive sample and a negative sample;

[0044] Among them, the determination process of the positive sample and the negative sample is:

[0045] A background bounding box is randomly generated around the target bounding box of each event frame or each image frame according to a uniform distribution algorithm;

[0046] An intersection-and-union ratio (IoU) of the background bounding box and the target bounding box is calculated, and an event frame or image frame whose IoU is greater than a first threshold is determined as a positive sample, and an event frame or image frame whose IoU is less than a second threshold is determined as a negative sample.

[0047] A target tracking system based on a pulse convolutional neural network, comprising:

[0048] A target video sequence acquisition module is used to acquire a target video sequence to be identified; the target video sequence is a video sequence based on an image frame or a video sequence based on an event frame; one of the image frames and one of the event frames both include a target area and a background area;

[0049] A classification result determination module, used for inputting the target video sequence to be identified into the pulse convolutional neural network model to determine the classification result of the target area and the classification result of the background area; the classification result of the target area is used to determine the driving trajectory of the target;

[0050] The pulse convolutional neural network model includes at least one trained first pulse convolutional neural network; the first pulse convolutional neural network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a flat layer, a first fully connected layer and a second fully connected layer connected in sequence; the model of neurons in the first pulse convolutional neural network is a current-based leakage integration and release model.

[0051] According to the specific embodiments provided by the present invention, the present invention discloses the following technical effects:

[0052] The present invention provides a target tracking method and system based on a pulse convolutional neural network; the target video sequence of the present invention can be a video sequence based on image frames or a video sequence based on event frames, which can accurately achieve the purpose of target tracking; the pulse convolutional neural network of the present invention comprises a first convolutional layer, a second convolutional layer, a third convolutional layer, a flat layer, a first fully connected layer and a second fully connected layer which are connected in sequence. Compared with a target recognition network with dozens or hundreds of layers, the network structure of the present invention is simple and the processing efficiency is higher. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0054] Figure 1 A process schematic diagram of a target tracking method based on a pulse convolutional neural network of the present invention;

[0055] Figure 2 This is a visualization result diagram of a portion of the DVSOT21 dataset sequence according to an embodiment of the present invention;

[0056] Figure 3 This is a schematic diagram of the frame regression principle of an embodiment of the present invention;

[0057] Figure 4 A flow chart of a target tracking method based on a pulse convolutional neural network according to the present invention;

[0058] Figure 5 This is a structural diagram of a target tracking system based on a pulse convolutional neural network in the present invention. DETAILED DESCRIPTION

[0059] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0060] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0061] Embodiment 1

[0062] In view of the fact that existing target tracking methods cannot accept both image frame-based video sequence inputs and event frame-based video sequence inputs without any preprocessing, an embodiment of the present invention provides a target tracking method based on a pulse convolutional neural network that can process frame-based or event frame-based video sequences, which can extract complex features and achieve relatively accurate and stable tracking. This is also the first directly trained pulse neural network tracking method. The method includes: offline training, online fine-tuning and updating, and adjusting the tracking results with a border regressor. First, the pulse neural network is trained offline using training samples. In the tracking stage, the sequence to be tested can be used to generate a positive and negative sample fine-tuning network, and corresponding to each sequence to be tested, a specific border regressor will be trained to modify the tracking results of each frame of the sequence to be tested to make it closer to the actual position of the target.

[0063] like Figure 1 As shown, an embodiment of the present invention provides a target tracking method based on a pulse convolutional neural network, which includes the following steps.

[0064] Step 1: Generate positive and negative training samples using the tracking sequences used for training.

[0065] The GOT-10k dataset and the DVSOT21 dataset are used as training sets for the embodiments of the present invention, wherein the GOT-10k dataset is a frame-based target tracking dataset released by the Chinese Academy of Sciences, and the DVSOT21 dataset is an event-based target tracking dataset proposed in the embodiments of the present invention. Figure 2 The visualization results of some sequences of the DVSOT21 dataset are given. The targets tracked from left to right are: car, star, football, cube and cat.

[0066] For the GOT-10k dataset, the embodiment of the present invention directly generates training samples on a given traditional image frame, whereas for the DVSOT21 dataset, the embodiment of the present invention needs to collect all events within every 30 ms as an event frame and generate training samples based on this.

[0067] The embodiment of the present invention needs to generate more positive and negative training samples through the given data set. Specifically, for all given training sequences, each frame (each event frame or each image frame) has a bounding box calibrating the position of the target. The embodiment of the present invention randomly generates a background bounding box based on the uniform distribution around the target bounding box of each frame, and the generated background bounding box with an intersection over union (IoU) greater than 0.7 with the target bounding box is used as a positive sample, and the IoU less than 0.5 is used as a negative sample.

[0068] The reason for distinguishing positive and negative samples is that the SNN used for tracking is used as a classification network in the present invention to identify whether the input sample is a target or a background. Among them, the positive sample with a high overlap rate with the ground-truth is regarded as the target, while the feature information contained in the negative sample is far away from the target and is regarded as the background.

[0069] Step 2: Use the generated positive and negative samples to train the corresponding spiking convolutional neural network (SNN).

[0070] The pulse convolutional neural network used in the embodiment of the present invention includes 3 convolutional layers, 1 flat layer and 2 fully connected layers.

[0071] The first layer is a convolutional layer, which uses 96 7×7 filters with a step size of 2. The input is a matrix of batch×C×107×107, where batch represents the number of batches input to the SNN at one time. For frame-based input, C=3, and the pixel values ​​can be directly input into the SNN as current after normalization, and the input samples can be enlarged or reduced to a size of 107×107. For event-based input, C=1, when the maximum horizontal and vertical coordinate values ​​of the input samples exceed 107, the horizontal and vertical coordinates of all input events can be reduced to the range of [0, 107). When the maximum horizontal and vertical coordinate values ​​of the input samples are less than 107, they are directly put into the SNN, and the pulses of the neurons corresponding to the coordinates that did not generate events are set to 0. Then, after a batch normalization and a maximum pooling, the filter size is 3×3 and the step size is 2, and the output is a binary pulse matrix of batch×96×25×25.

[0072] The second layer is a convolutional layer, which uses 256 5×5 filters with a step size of 2 and an input binary impulse matrix of batch×96×25×25. After a maximum pooling, where the filter size is 3×3 and the step size is 2, the output is a binary impulse matrix of batch×256×11×11.

[0073] The third layer is a convolutional layer, which uses 512 3×3 filters with a step size of 1 and an input of a binary impulse matrix of batch×256×11×11. This layer does not require pooling, and the output is a binary impulse matrix of batch×512×3×3.

[0074] The fourth layer is a flattening layer, which stretches the batch×512×3×3 binary pulse matrix input into a batch×4608 binary pulse vector output. The number of output neurons in this layer is 4608.

[0075] The fifth layer is a fully connected layer. Its input is a binary pulse vector of batch×4608 and its output is a binary pulse vector of batch×512. The number of output neurons in this layer is 512.

[0076] The sixth layer is a fully connected layer. Its input is a binary pulse vector of batch×512, and its output is a binary pulse vector of batch×2. The number of output neurons in this layer is 2, one of which represents that the input sample is the background (negative sample) and the other represents that the input sample is the target (positive sample).

[0077] Two SNNs are trained based on the input image frame-based video sequence and event frame-based video sequence. Except for the different inputs, the network structure, neuron model and learning algorithm of the SNN are the same. The neuron model used in the embodiment of the present invention is a current based leaky integrate and fire (C-LIF) model, and its dynamic equation is as follows:

[0078] V t =L t -E t

[0079]

[0080]

[0081]

[0082] Among them, V t represents the membrane voltage of the neuron at the tth moment (in the absence of input, the membrane voltage of the spiking neuron decreases with the passage of time), L t represents the membrane voltage calculated by the kinetic equation of the Leaky Integrate and Fire (LIF) model at time t, E t represents the reset item at time t, E t The difference between the C-LIF model and the LIF model is that each output pulse will suppress the current membrane voltage. N means there are N input neurons, w i represents the synaptic weight of the i-th input neuron, represents the moment when the i-th input neuron in the previous layer emits the j-th output pulse, It indicates the time when the sth input neuron in the current layer emits the jth output pulse. Each pulse at time t contributes a postsynaptic potentia (PSP), whose shape is determined by a double exponential kernel function Determine; K represents the standardized postsynaptic potential kernel. V0 is the normalization factor that makes the maximum value of the kernel function 1. α m ,α s represents the learnable time decay factor, and θ represents the threshold of the neuron. For the C-LIF model, whenever the membrane voltage exceeds the threshold, a pulse is fired, and then the membrane voltage is inhibited.

[0083] The embodiment of the present invention uses the Spatio-Temporal Credit Assignment (STCA) learning algorithm, which is a supervised learning algorithm for multi-pulse output of training deep SNN based on the C-LIF model. The feedforward calculation of the network is as follows:

[0084]

[0085]

[0086]

[0087]

[0088]

[0089]

[0090] Among them, the subscript i, superscript k and t represent the state of the i-th neuron in the k-th layer at time t.

[0091] Represents the superposition of the input response and output response of the i-th neuron in the k-th layer at the t-th time; represents the single-core input response of the i-th neuron in the k-th layer at the t-th time, corresponding to the first exponential kernel of K in the dynamic equation; represents the single-core input response of the i-th neuron in the k-th layer at the t-th time, corresponding to the second exponential kernel of K in the dynamic equation; Represents the inhibition produced by the output of the i-th neuron in the k-th layer at the t-th time; represents the weighted sum of the pulse inputs of the i-th neuron at the k-th layer at the t-th time; Indicates whether the i-th neuron in the k-th layer at the t-1th time has a pulse generated, 1 indicates that a pulse is generated, and 0 indicates that no pulse is generated. represents the synaptic weight.

[0092] L(k) represents the number of neurons in the kth layer. MS represents the response to the input pulse, and E represents the inhibition of the output pulse, both of which are affected by the state of the previous moment.

[0093] I represents the weighted sum of the input pulses, and O represents whether a pulse is generated (1 means there is a pulse, and 0 means there is no pulse). m , α s represents the learnable time decay factor, and θ represents the threshold of the neuron. The feedback calculation of the network is based on the principle of gradient descent, similar to the backpropagation through time (BPTT) learning algorithm.

[0094] The loss function of the network is related to the IoU between the training sample box and the target bounding box. Assume that the IoU between the training sample box and the target bounding box is O s , s represents a positive sample or a negative sample, then the contribution calculation formula of a sample is as follows:

[0095]

[0096] Based on this, the loss function of the network is defined as follows:

[0097]

[0098] Among them, V max Represents the maximum voltage of the output neuron at all time steps, θ+R s Represents the expected output voltage for one sample.

[0099] Step 3: Fine-tune the trained SNN using the test sequence, and at the same time train a bounding box regressor specific to the test sequence.

[0100] When a sequence to be tested is given, and only the position of the target in the first frame is specified, that is, only the target bounding box of the first frame is given, the subsequent target bounding boxes need to be predicted by the pulse convolutional neural network of the present invention. In order to improve the accuracy of the tracking results, the present invention will fine-tune the trained SNN according to the information of the first frame of the sequence to be tested. The embodiment of the present invention randomly generates 500 positive sample frames near the target bounding box of the first frame according to the normal distribution, and randomly generates 1000 negative sample frames near the target bounding box of the first frame according to the uniform distribution. Finally, 1000 negative sample frames are randomly generated within the entire image range of the first frame, and these 2500 sample frames are used to fine-tune the SNN.

[0101] In addition, in order to further improve the tracking accuracy, the present invention also uses the information of the first frame of the sequence to be tested to train a bounding box regressor, which is essentially a ridge regression. Figure 3 As shown in the figure, the P bounding box is the target position predicted by SNN, and the G bounding box is the actual location of the target. The purpose of bounding box regression is to modify P to obtain the G' bounding box so that the predicted position is closer to the actual location of the target.

[0102] The present invention randomly generates 1000 training sample frames whose IoU with the target bounding box is greater than 0.6 based on uniform distribution around the target bounding box in the first frame, and sends these sample frames to SNN to extract features, and then uses the features output by the third convolutional layer to train the frame regressor.

[0103] The trained SNN is trained based on the features of the targets in the training set, which may be different from the target types or features in the sequence to be tested. In order to improve the tracking accuracy, the known information on the sequence to be tested can be used, that is, positive and negative samples are generated on the first frame to fine-tune the SNN, so that the weights of the SNN contain feature information specific to the sequence to be tested.

[0104] As in the training phase, the positive and negative samples generated in the first frame are used to adjust the weights of the SNN. The Adam optimizer is also used, but the learning rate at this time is only half of that in the training phase.

[0105] The essence of the bounding box regressor in this application is ridge regression, which was first used in object detection to adjust the predicted object bounding box to make it closer to the ground-truth.

[0106] Step 4: Use the tuned SNN to track the result and use the trained bounding box regressor to modify and get the final result.

[0107] For each test sequence, the present invention starts predicting the position of the target in the second frame. The specific tracking process is as follows: the present invention randomly generates 256 target candidate frames near the target bounding box predicted in the previous frame according to the normal distribution, and puts these candidate frames into the currently predicted frame, extracts the image information in the frame (frame input: extract the image in the frame, event input: extract all events in the frame) and inputs the SNN for evaluation. Each target candidate frame will generate a corresponding score, that is, the voltage of the target neuron of the sixth fully connected layer. The present invention selects the target candidate frame corresponding to the highest voltage from the 256 voltages as the target bounding box predicted for the frame.

[0108] Finally, the present invention utilizes the trained bounding box regressor to modify the currently predicted target bounding box to make it closer to the actual location of the target.

[0109] Embodiment 2

[0110] In order to realize target tracking of a pulse convolutional neural network that can process both frame-based input and event-based input, an embodiment of the present invention provides a target tracking method based on a pulse convolutional neural network, such as Figure 4 As shown, including:

[0111] Step 100: Acquire a target video sequence to be identified; the target video sequence is a video sequence based on an image frame or a video sequence based on an event frame; one of the image frames and one of the event frames both include a target area and a background area.

[0112] Step 200: Input the target video sequence to be identified into the pulse convolutional neural network model to determine the classification result of the target area and the classification result of the background area; the classification result of the target area is used to determine the driving trajectory of the target.

[0113] The pulse convolutional neural network model includes at least one trained first pulse convolutional neural network; the first pulse convolutional neural network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a flat layer, a first fully connected layer and a second fully connected layer connected in sequence; the model of neurons in the first pulse convolutional neural network is a current-based leakage integration and release model.

[0114] Among them, when the target video sequence includes a video sequence based on image frames and a video sequence based on event frames, the pulse convolutional neural network model also includes a trained second pulse convolutional neural network with the same structure as the first pulse convolutional neural network; the first pulse convolutional neural network is used to determine the classification result of the target area and the classification result of the background area according to the video sequence based on image frames; the second pulse convolutional neural network is used to determine the classification result of the target area and the classification result of the background area according to the video sequence based on event frames.

[0115] In one example, step 200 specifically includes:

[0116] Step A: preprocessing the target video sequence to be identified.

[0117] Step B: fine-tuning the parameters of the pulse convolutional neural network model according to the target area and background area of ​​the first frame of the preprocessed target video sequence; the first frame is an image frame or an event frame.

[0118] Step C: Input the preprocessed target video sequence into the fine-tuned pulse convolutional neural network model to determine the classification result of the target area and the classification result of the background area.

[0119] Furthermore, step A specifically includes:

[0120] When the target video sequence is a video sequence based on image frames, the target video sequence to be identified is normalized to obtain a pre-processed target video sequence.

[0121] When the target video sequence is a video sequence based on event frames, the target video sequence to be identified is subjected to horizontal and vertical coordinate reduction processing to obtain a preprocessed target video sequence.

[0122] Furthermore, step B specifically includes:

[0123] According to the first frame of the preprocessed target video sequence, a fine-tuning sample set is obtained; the fine-tuning sample set includes a plurality of positive sample frames randomly generated near the target bounding box of the first frame using a normal distribution algorithm, a plurality of first negative sample frames randomly generated near the target bounding box of the first frame according to a uniform distribution algorithm, and a plurality of second negative sample frames randomly generated within the entire image range of the first frame according to a uniform distribution algorithm; wherein an intersection-over-union ratio of the positive sample frame to the target bounding box is greater than a first threshold; and an intersection-over-union ratio of the negative sample frame to the target bounding box is less than a second threshold.

[0124] The parameters of the pulse convolutional neural network model are fine-tuned according to the fine-tuning sample set.

[0125] Furthermore, step C specifically includes:

[0126] A regressor sample set is obtained according to the first frame of the preprocessed target video sequence; the regressor sample set includes a plurality of training sample frames randomly generated around the target bounding box of the first frame according to a uniform distribution algorithm; an intersection-over-union ratio of the training sample frame to the target bounding box is greater than a third threshold.

[0127] The regressor sample set is input into the fine-tuned pulse convolutional neural network model so that the features of the third convolutional layer output in the fine-tuned pulse convolutional neural network model are used to train the bounding box regressor to obtain a trained bounding box regressor.

[0128] The preprocessed target video sequence is input into the fine-tuned pulse convolutional neural network model to obtain the classification results of the target area and the classification results of the background area; the marking features output by the third convolutional layer in the fine-tuned pulse convolutional neural network model are input into the trained border regressor to obtain the correction result of the target area; the marking features are the features output by the third convolutional layer after the preprocessed target video sequence is input into the fine-tuned pulse convolutional neural network model.

[0129] The classification result of the target area is classified according to the correction result of the target area to obtain a corrected classification result of the target area.

[0130] In one example, the first convolutional layer includes 96 7×7 filters; the stride of the first convolutional layer is 2; the input of the first convolutional layer is a matrix of batch×C×107×107, and the output of the first convolutional layer is a binary pulse matrix of batch×96×25×25; wherein batch represents the number of batches input into the first pulse convolutional neural network or the second pulse convolutional neural network at one time; when the target video sequence is a video sequence based on image frames, C=3; when the target video sequence is a video sequence based on event frames, C=1.

[0131] The second convolutional layer includes 256 5×5 filters; the step size of the second convolutional layer is 2; the input of the second convolutional layer is a binary pulse matrix of batch×96×25×25, and the output of the second convolutional layer is a binary pulse matrix of batch×256×11×11.

[0132] The third convolutional layer includes 512 3×3 filters; the step size of the third convolutional layer is 1; the input of the third convolutional layer is a binary pulse matrix of batch×96×25×25, and the output of the third convolutional layer is a binary pulse matrix of batch×512×3×3.

[0133] The flattening layer is used to stretch the input batch×512×3×3 binary pulse matrix into a batch×4608 binary pulse vector.

[0134] The input of the first fully connected layer is a binary pulse vector of batch×4608, and the output of the first fully connected layer is a binary pulse vector of batch×512.

[0135] The input of the second fully connected layer is a binary pulse vector of batch×512, and the output of the second fully connected layer is a binary pulse vector of batch×2; the number of neurons output by the second fully connected layer is 2, wherein one neuron is used to output the classification result of the target area, and the other neuron is used to output the classification result of the background area.

[0136] In one example, the kinetic equation of the current-based leakage integration release model is:

[0137] V t =L t -E t

[0138]

[0139]

[0140]

[0141] Among them, V t represents the membrane voltage of the neuron, L t represents the membrane voltage calculated from the kinetic form of the leak-integration-spark model, E t represents the reset item; N means there are N input neurons, w i represents the synaptic weight of the i-th input neuron, represents the time when the i-th input neuron emits the j-th input pulse, represents the time when the sth input neuron emits the jth output pulse; V0 represents the normalization factor; α m , α s Both represent learnable time decay factors, and θ represents the threshold of the neuron.

[0142] In one example, the first trained pulse convolutional neural network and the second trained pulse convolutional neural network are both determined by training the pulse convolutional neural network using a training set and a spatiotemporal credit allocation learning algorithm;

[0143] The training set includes multiple sample pairs; each of the sample pairs includes a positive sample and a negative sample.

[0144] Among them, the determination process of the positive sample and the negative sample is:

[0145] A background bounding box is randomly generated around the target bounding box in each event frame or each image frame according to a uniform distribution algorithm.

[0146] An intersection-and-union ratio (IoU) of the background bounding box and the target bounding box is calculated, and an event frame or image frame whose IoU is greater than a first threshold is determined as a positive sample, and an event frame or image frame whose IoU is less than a second threshold is determined as a negative sample.

[0147] To achieve the above objectives, the present invention also provides a target tracking system based on a pulse convolutional neural network. Figure 5 As shown, including:

[0148] The target video sequence acquisition module 300 is used to acquire a target video sequence to be identified; the target video sequence is a video sequence based on an image frame or a video sequence based on an event frame; one of the image frames and one of the event frames both include a target area and a background area;

[0149] The classification result determination module 400 is used to input the target video sequence to be identified into the pulse convolutional neural network model to determine the classification result of the target area and the classification result of the background area; the classification result of the target area is used to determine the driving trajectory of the target;

[0150] The pulse convolutional neural network model includes at least one trained first pulse convolutional neural network; the first pulse convolutional neural network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a flat layer, a first fully connected layer and a second fully connected layer connected in sequence; the model of neurons in the first pulse convolutional neural network is a current-based leakage integration and release model.

[0151] The present invention provides a target tracking method and system based on a pulse neural network, which has the following advantages:

[0152] 1. The network structure is simple, low power consumption, and low latency. Compared with the target recognition network with dozens or hundreds of layers, the network structure of the present invention is simple, with only six layers. Moreover, compared with the traditional convolutional neural network with the same structure, the implementation of the pulse convolutional neural network on neuromorphic hardware has the characteristics of low power consumption and low latency.

[0153] 2. It can process image frame-based video sequences and event frame-based video sequences with high precision; existing video sequence tracking technologies can only process image frame-based video sequences or event frame-based video sequences, while the present invention is a tracking algorithm based on a pulse convolutional neural network, which can regard the input event frame-based video sequence as a pulse sequence, and for image frame-based video sequences, the pixel values ​​can also be directly input into the SNN after normalization. Therefore, the present invention can accept the input of both image frame-based video sequences and event frame-based video sequences without any preprocessing, and has high tracking accuracy. The tracking results of the present invention and an existing conversion SNN network on the frame-based GOT-10k data set are compared as shown in Table 1.

[0154] Table 1 Comparison of tracking results of GOT-10k dataset

[0155] Tracker Comparative Example 1 The present invention AO 0.314 0.320 <![CDATA[SR 0藀5 ]]> 0.327 0.346

[0156] Comparative Example 1: Yihao Luo, Min Xu, Caihong Yuan, Xiang Cao, Liangqi Zhang, Yan Xu, Tianjiang Wang, and Qi Feng. Siamsnn: Siamese spiking neural networks for energy-efficient object tracking. In International Conference on Artificial Neural Networks, pages 182–194. Springer, 2021.

[0157] Among them, AO represents the average overlap rate between the predicted target bounding box and the real bounding box in the test sequence, and SR 0藀5 The percentage of successfully tracked frames with an overlap rate greater than 0.5. The present invention is superior to several existing trackers based on traditional correlation filtering, traditional deep neural network, and event-based trackers.

[0158] The comparison of tracking results on the DVSOT21 dataset is shown in Table 2.

[0159] Table 2 Comparison of tracking results of DVSOT21 dataset

[0160] Tracker Comparative Example 2 Comparative Example 3 Comparative Example 4 Comparative Example 5 The present invention AO 0.628 0.772 0.431 0.223 0.748 <![CDATA[SR 0藀5 ]]> 0.736 0.925 0.850 0.384 0.979

[0161] Comparative Example 2: David S Bolme, J Ross Beveridge, Bruce A Draper, and Yui ManLui.Visual object tracking using adaptive correlation filters.In 2010IEEEcomputer society conference on computer vision and pattern recognition, pages2544–2550.IEEE, 2010.

[0162] 3Jo ao F Henriques, Rui Caseiro, Pedro Martins, and Jorge Batista. High-speed tracking with kernelized correlation filters. IEEE transactions onpattern analysis and machine in-telligence, 37(3):583–596, 2014.

[0163] Comparative Example 4: David Held, Sebastian Thrun, and Silvio Savarese. Learning totrack at 100fps with deep regression networks. In European conference on computer vision, pages 749–765. Springer, 2016.

[0164] Comparative Example 5: TD Brndli Christian and Luca Longinotti.jaer open source project, 2014.

[0165] 3. The first directly trained SNN tracking network: This invention is the first target tracking network directly trained with SNN, which can directly accept the event stream output by the event camera without converting it into frame input first. Compared with the current event-based tracking algorithm, which is easily affected by noise events and deviates from the tracking target, this invention can extract more complex and robust target features and achieve more stable tracking.

[0166] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.

[0167] The principles and implementation methods of the present invention are described in this article using specific examples. The description of the above embodiments is only used to help understand the method and core idea of ​​the present invention. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting the present invention.

Claims

1. A target tracking method based on a pulse convolutional neural network, characterized in that: include: Acquire a target video sequence to be identified; the target video sequence is a video sequence based on image frames or a video sequence based on event frames; One of the image frames and one of the event frames both include a target area and a background area; Inputting the target video sequence to be identified into a pulse convolutional neural network model to determine the classification result of the target area and the classification result of the background area; The classification result of the target area is used to determine the driving trajectory of the target; The pulse convolutional neural network model includes at least one trained first pulse convolutional neural network; the first pulse convolutional neural network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a flat layer, a first fully connected layer and a second fully connected layer connected in sequence; the model of the neurons in the first pulse convolutional neural network is a leakage integrated release model based on current; Before inputting the target video sequence to be identified into the pulse convolutional neural network model, the method further includes: preprocessing the target video sequence to be identified; The preprocessing of the target video sequence to be identified specifically includes: When the target video sequence is a video sequence based on image frames, normalizing the target video sequence to be identified to obtain a preprocessed target video sequence; When the target video sequence is a video sequence based on event frames, performing horizontal and vertical coordinate reduction processing on the target video sequence to be identified to obtain a preprocessed target video sequence; The parameters of the pulse convolutional neural network model are fine-tuned according to the target area and the background area of ​​the first frame of the preprocessed target video sequence; the first frame is an image frame or an event frame.

2. The target tracking method based on pulse convolutional neural network according to claim 1 is characterized in that: When the target video sequence includes a video sequence based on image frames and a video sequence based on event frames, the pulse convolutional neural network model also includes a trained second pulse convolutional neural network with the same structure as the first pulse convolutional neural network; The first pulse convolutional neural network is used to determine the classification result of the target area and the classification result of the background area according to the video sequence based on the image frame; The second pulse convolutional neural network is used to determine the classification result of the target area and the classification result of the background area according to the event frame-based video sequence.

3. The target tracking method based on pulse convolutional neural network according to claim 1 is characterized in that: The step of inputting the target video sequence to be identified into a pulse convolutional neural network model to determine the classification result of the target area and the classification result of the background area specifically includes: The preprocessed target video sequence is input into the fine-tuned pulse convolutional neural network model to determine the classification result of the target area and the classification result of the background area.

4. The target tracking method based on pulse convolutional neural network according to claim 1 is characterized in that: The parameters of the spike convolutional neural network model are fine-tuned according to the target area and background area of ​​the first frame of the preprocessed target video sequence, including: According to the first frame of the preprocessed target video sequence, a fine-tuning sample set is obtained; the fine-tuning sample set includes a plurality of positive sample frames randomly generated near the target bounding box of the first frame using a normal distribution algorithm, a plurality of first negative sample frames randomly generated near the target bounding box of the first frame using a uniform distribution algorithm, and a plurality of second negative sample frames randomly generated within the entire image range of the first frame using a uniform distribution algorithm; wherein an intersection-over-union ratio of the positive sample frame to the target bounding box is greater than a first threshold; and an intersection-over-union ratio of the negative sample frame to the target bounding box is less than a second threshold; The parameters of the pulse convolutional neural network model are fine-tuned according to the fine-tuning sample set.

5. The target tracking method based on pulse convolutional neural network according to claim 3 is characterized in that: The preprocessed target video sequence is input into the fine-tuned pulse convolutional neural network model to determine the classification result of the target area and the classification result of the background area, specifically including: According to the first frame of the preprocessed target video sequence, a regressor sample set is obtained; the regressor sample set includes a plurality of training sample frames randomly generated around the target bounding box of the first frame according to a uniform distribution algorithm; the intersection and union ratio of the training sample frame and the target bounding box is greater than a third threshold; Inputting the regressor sample set into the fine-tuned spiking convolutional neural network model so that the features of the third convolutional layer output in the fine-tuned spiking convolutional neural network model are used to train a bounding box regressor to obtain a trained bounding box regressor; Inputting the preprocessed target video sequence into the fine-tuned pulse convolutional neural network model to obtain the classification result of the target area and the classification result of the background area; inputting the marking feature output by the third convolutional layer in the fine-tuned pulse convolutional neural network model into the trained border regressor to obtain the correction result of the target area; the marking feature is the feature output by the third convolutional layer after the preprocessed target video sequence is input into the fine-tuned pulse convolutional neural network model; The classification result of the target area is classified according to the correction result of the target area to obtain a corrected classification result of the target area.

6. The target tracking method based on pulse convolutional neural network according to claim 2 is characterized in that: The first convolutional layer includes 96 7×7 filters; the stride of the first convolutional layer is 2; the input of the first convolutional layer is a matrix of batch×C×107×107, and the output of the first convolutional layer is a binary impulse matrix of batch×96×25×25; wherein batch represents the number of batches input into the first impulse convolutional neural network or the second impulse convolutional neural network at one time; when the target video sequence is a video sequence based on image frames, C=3; when the target video sequence is a video sequence based on event frames, C=1; The second convolutional layer includes 256 5×5 filters; the stride of the second convolutional layer is 2; the input of the second convolutional layer is a binary pulse matrix of batch×96×25×25, and the output of the second convolutional layer is a binary pulse matrix of batch×256×11×11; The third convolutional layer includes 512 3×3 filters; the step size of the third convolutional layer is 1; the input of the third convolutional layer is a binary pulse matrix of batch×96×25×25, and the output of the third convolutional layer is a binary pulse matrix of batch×512×3×3; The flattening layer is used to stretch the input batch×512×3×3 binary pulse matrix into a batch×4608 binary pulse vector; The input of the first fully connected layer is a binary pulse vector of batch×4608, and the output of the first fully connected layer is a binary pulse vector of batch×512; The input of the second fully connected layer is a binary pulse vector of batch×512, and the output of the second fully connected layer is a binary pulse vector of batch×2; the number of neurons output by the second fully connected layer is 2, wherein one neuron is used to output the classification result of the target area, and the other neuron is used to output the classification result of the background area.

7. The target tracking method based on pulse convolutional neural network according to claim 2 is characterized in that: The kinetic equation of the current-based leak-integrated-release model is: V t =L t -E t Among them, V t represents the membrane voltage of the neuron at time t, L t represents the membrane voltage calculated by the kinetic equation of the leak-integration-spiking model at time t, E t represents the reset item at time t; N represents N input neurons, w i represents the synaptic weight of the i-th input neuron, represents the moment when the i-th input neuron in the previous layer emits the j-th output pulse, Indicates when The time when the sth input neuron in the front layer emits the jth output pulse; V0 represents the normalization factor; α m , α s Both represent learnable time decay factors, θ represents the threshold of the neuron; K represents the standardized postsynaptic potential kernel.

8. The target tracking method based on pulse convolutional neural network according to claim 2 is characterized in that: The first trained pulse convolutional neural network and the second trained pulse convolutional neural network are both determined after training the pulse convolutional neural network using a training set and a spatiotemporal credit allocation learning algorithm; The training set includes a plurality of sample pairs; each of the sample pairs includes a positive sample and a negative sample; Among them, the determination process of the positive sample and the negative sample is: A background bounding box is randomly generated around the target bounding box of each event frame or each image frame according to a uniform distribution algorithm; An intersection-and-union ratio (IoU) of the background bounding box and the target bounding box is calculated, and an event frame or image frame whose IoU is greater than a first threshold is determined as a positive sample, and an event frame or image frame whose IoU is less than a second threshold is determined as a negative sample.

9. A target tracking system based on a pulse convolutional neural network, characterized in that: include: A target video sequence acquisition module is used to acquire a target video sequence to be identified; The target video sequence is a video sequence based on image frames or a video sequence based on event frames; One of the image frames and one of the event frames both include a target area and a background area; A classification result determination module, used for inputting the target video sequence to be identified into a pulse convolutional neural network model to determine the classification result of the target area and the classification result of the background area; The classification result of the target area is used to determine the driving trajectory of the target; Before inputting the target video sequence to be identified into the pulse convolutional neural network model, the method further includes: preprocessing the target video sequence to be identified; The preprocessing of the target video sequence to be identified specifically includes: When the target video sequence is a video sequence based on image frames, normalizing the target video sequence to be identified to obtain a preprocessed target video sequence; When the target video sequence is a video sequence based on event frames, performing horizontal and vertical coordinate reduction processing on the target video sequence to be identified to obtain a preprocessed target video sequence; Fine-tuning the parameters of the spike convolutional neural network model according to the target area and the background area of ​​the first frame of the preprocessed target video sequence; the first frame is an image frame or an event frame; The pulse convolutional neural network model includes at least one trained first pulse convolutional neural network; the first pulse convolutional neural network includes a first convolutional layer, a second convolutional layer, a third convolutional layer, a flat layer, a first fully connected layer and a second fully connected layer connected in sequence; the model of neurons in the first pulse convolutional neural network is a current-based leakage integration and release model.

Citation Information

Patent Citations

  • V2V video face recognition method based on common feature subspace

    CN113033345A

  • Gesture recognition method and recognition system

    CN113205048A