Method and device for tracking single target on water surface of unmanned ship based on spiking neural network

By constructing an end-to-end target tracking model of pulsed neural network, the application problem of pulsed neural network in single-target tracking of unmanned boats is solved, low-power real-time target tracking is achieved, and the problem of indirection of direct training of pulsed neural networks is overcome.

CN120298450APending Publication Date: 2025-07-11YICHANG TESTING TECHNIQUE RESEARCH INSTITUTE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311493460.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-11-09
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

The prior art is difficult to apply pulsed neural networks to single-target tracking scenarios on unmanned boats surfaces, and there are difficulties in direct training of pulsed neural networks, which leads to difficulties in applying deep neural networks in resource-constrained unmanned boat systems.

Method used

An end-to-end target tracking model based on pulsed neural network is constructed. By modifying the SiamFC model, the parameters of the batch sample normalization layer are merged into the convolution layer weights of the previous layer, the bias of the last layer of the backbone convolutional neural network is deleted, the model is trained using visible and infrared light images, and the response map is generated through two-stage periodic coding and similarity estimation methods.

Benefits of technology

It realizes low-power real-time target tracking in resource-constrained unmanned boat systems, solves the application problem of pulsed neural networks in single-target tracking of unmanned boats surfaces, reduces inference time and supports real-time inference.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298450A_ABST
    Figure CN120298450A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned ship water surface single target tracking method and device based on a spiking neural network. The method comprises the following steps: constructing an end-to-end target tracking model based on the spiking neural network; acquiring a template image of a target needing to be tracked, and extracting a pulse feature sequence of the template image; obtaining a current frame image, and extracting a pulse feature sequence of the current frame image; determining a similarity response diagram of the pulse feature sequences of the template image and the current frame image; and determining a target tracking result of the current frame image based on the similarity response diagram. According to the method, high-precision detection and identification of the underwater target are realized. According to the invention, the spiking neural network is applied to the unmanned ship water surface single target tracking task, and by means of the characteristic of low power consumption in the calculation reasoning process, the difficulty of applying the deep neural network in the unmanned ship system with limited resources is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and unmanned surface vessels, and particularly relates to a method and device for single-object tracking on the water surface of an unmanned vessel based on a spiking neural network. Background Art

[0002] Unmanned surface vessels can perform tasks in various complex water area environments, and do not require operators to drive on the vessel, ensuring the safety of personnel. Unmanned vessels play important roles in both civilian and military applications. With the increasing demand for the application of unmanned vessels in China, the requirements for performance indicators such as environmental perception, autonomous obstacle avoidance, and operating energy consumption are gradually increasing. Visible light cameras and infrared cameras have the advantages of low cost, low power consumption, and easy deployment, and have become the mainstream information acquisition sensors for unmanned surface vessels.

[0003] Object tracking is a key technology required for the environmental perception of unmanned vessels, which can locate and track a certain object in a continuously acquired image sequence, and provide target information for subsequent autonomous obstacle avoidance and navigation control. In early research, object tracking was mainly based on correlation filtering algorithms, using traditional features for object matching. With the maturity of big data and deep learning related technologies, object tracking algorithms based on siamese neural networks have achieved better performance in many unmanned application tasks, represented by models such as SiamFC, SiamRPN, and Siam R-CNN. However, these models usually need to be deployed on high-performance graphics cards for operation, and the power consumption generated far exceeds that of ordinary computing devices. In the application scenarios of small unmanned vessels, the computing power and power consumption of computing devices are limited, and it is difficult to carry high-performance graphics cards for long-term operation.

[0004] As the "third generation of neural networks", spiking neural networks simulate the characteristics of neurons in the biological nervous system to transmit information through spike signals, and are a popular research topic for realizing brain-like intelligence. Spiking neural networks transmit information with discrete binary "0, 1" signals. Neurons accumulate the received spike signals, and will output a spike signal only when the accumulated potential exceeds a preset threshold. This information transmission method enables spiking neural networks not to require multiplication during the inference process, and the mode of electrical signal superposition can achieve low-power operation when designing chips.

[0005] Currently, although there are already some methods to apply spiking neural networks to image classification and object detection, they have not been applied to the scenario of single-object tracking on the water surface of unmanned vessels. In addition, there are difficulties such as non-differentiability in directly training spiking neural networks, and the implementation accuracy of some unsupervised training methods is difficult to meet the requirements. Summary of the Invention

[0006] In view of this, the present invention provides a method and device for single-object tracking on the water surface of an unmanned boat based on a spiking neural network, which can solve the technical problems of applying a spiking neural network to the single-object tracking scenario on the water surface of an unmanned boat and the non-differentiability of directly training a spiking neural network.

[0007] To solve the above technical problems, the present invention is implemented as follows.

[0008] A method for single-object tracking on the water surface of an unmanned boat based on a spiking neural network, the method comprising:

[0009] Step S1: Construct an end-to-end target tracking model based on a spiking neural network;

[0010] Step S2: Obtain a template image of the target to be tracked, and extract the spiking feature sequence of the template image;

[0011] Step S3: Obtain the current frame image, and extract the spiking feature sequence of the current frame image;

[0012] Step S4: Determine the similarity response map of the spiking feature sequences of the template image and the current frame image;

[0013] Step S5: Based on the similarity response map, determine the target tracking result of the current frame image;

[0014] Step S6: Read the next frame image from the video stream as the current frame image, and enter Step S3; if there is no next frame image, the method ends.

[0015] Preferably, the Step S1: Construct an end-to-end target tracking model based on a spiking neural network, includes:

[0016] Step S11: Modify the SiamFC model, including: merging all the parameters of the batch normalization layer (BatchNormalization, BN) of the SiamFC model into the convolutional layer weights of the previous layer; the calculation formula and parameter merging formula of the batch normalization layer are:

[0017]

[0018]

[0019]

[0020] where x is the input vector of the BN layer, BN[x] is the output vector of the BN layer, μ and σ are the mean and variance of the training data respectively, and β and γ are variables obtained by training through gradient descent; is the convolutional kernel weight of the l-th layer after modification, is the weight coefficient of the l-th layer BN layer, is the variance of the training data of the l-th layer, is the weight of the convolutional kernel of the l-th layer; is the modified bias of the l-th layer convolutional layer, is the bias of the l-th layer convolutional layer, is the mean of the training data of the l-th layer, is the constant coefficient of the l-th layer BN layer;

[0021] Delete the bias of the last layer of the backbone convolutional neural network of the SiamFC model;

[0022] Step S12: Use the first training set to perform several rounds of iterative training on the modified SiamFC model, and then use the second training set including visible light images and infrared light images to train the modified SiamFC model to obtain the trained modified SiamFC model;

[0023] Step S13: Convert the trained modified SiamFC model into a spiking neural network model;

[0024] Step S14: Generate a potential threshold through a two-stage periodic coding method, and use the potential threshold as the potential threshold of the spiking convolutional network, including:

[0025] The two-stage periodic coding refers to encoding the spiking feature sequence in a two-stage threshold manner, divided into two potential thresholds; based on the two potential thresholds, the potential threshold of the spiking neuron is periodically changed. The formula for calculating the potential threshold is:

[0026]

[0027] where p is the threshold change period, t is the current pulse simulation time, and ε is the activation threshold;

[0028] Step S15: Establish a similarity calculation formula for the spiking neural network model, where:

[0029] Similarity refers to the similarity between the target template and the spiking feature sequence of the current frame search area, including temporal similarity and potential similarity;

[0030] The temporal similarity considers the direct temporal correlation and window temporal correlation of the spiking feature sequence. The direct temporal correlation directly calculates the convolution result of the spiking feature sequence at each time step. The window temporal correlation takes into account that the correlation of spiking signals with a farther time interval is lower, and determines the temporal similarity according to the convolution result of the signal values of the two spiking sequences within the sliding time window 2τ. The formula for calculating the temporal similarity is as follows:

[0031]

[0032] Among them, z and x represent the target template image and the current frame search area image respectively, represents the pulse convolution network, τ is the time window parameter, T is the total simulation step of the spiking neural network, and k is the weight set in the experiment;

[0033] The potential similarity considers the overall potential correlation of the pulse feature sequence and calculates the convolution result of the synthesis of the pulse signals in all simulation time steps. The calculation formula is as follows:

[0034]

[0035] Among them, is the bias.

[0036] Step S16: Generate the response map of the spiking neural network model.

[0037] Preferably, the step S13: converting the trained modified SiamFC model into a spiking neural network model includes:

[0038] Calculating the response value based on the pixel value of each pixel point of the image as the input of the spiking neural network model and the convolution kernel parameter of the first layer of the backbone convolutional neural network of the trained modified SiamFC model, and taking the response value as the fixed potential value input of the first layer of the spiking neural network of the spiking neural network model; the calculation formula of the response value is:

[0039]

[0040] Among them, is the output response value of the first layer of the spiking neural network, V th is the potential threshold, is the weight of the first layer of the spiking neural network, x j is the input value of the first layer of the spiking neural network, is the bias of the first layer of the spiking neural network;

[0041] For the backbone convolutional neural network of the trained modified SiamFC model, from the second layer to the last layer, the parameter normalization method is used layer by layer to convert the backbone convolutional neural network weights into the weights of the pulse convolution network of the spiking neural network model. The conversion formula is:

[0042]

[0043] Among them, the normalization factor λ l is the 99.9% largest real number in the l-th layer activation value vector, w l 、bl They are the weights and biases of the l-th convolutional layer respectively, They are the weights and biases of the converted spiking convolutional network respectively;

[0044] The converted backbone convolutional neural network of SiamFC is the spiking convolutional network, and its output features serve as the spiking feature sequence of the spiking neural network model.

[0045] Preferably, step S16: generating the response map of the spiking neural network model includes:

[0046] Based on the temporal similarity and potential similarity between the target template and the current frame search area, linearly weighted fusion of the temporal similarity and potential similarity is performed to obtain a two-dimensional matrix of spiking feature sequence similarity. The calculation formula is

[0047] f s (z, x) = f P (z, x) + λf T (z, x),

[0048] where λ is the weight set by experiment;

[0049] Then traverse the two-dimensional similarity matrix generated by each batch in the inference result of the current spiking neural network model, and select the matrix with the maximum similarity value; subtract the minimum element value in the matrix from each element in the matrix, and divide by the sum of all elements in the matrix for normalization; finally, linearly fuse the hanning window matrix with the similarity matrix to punish the points at the coordinate positions far from the target, and the obtained fusion result is used as the response map of the image input to the spiking neural network model. The calculation formula is as follows:

[0050] response map = (1 - α)·f(z, x) + α·f hanning

[0051] where f is the normalized similarity matrix, f hanning is the hanning window matrix, α is the weight set by experiment, and f(z, x) is the similarity matrix of normalized z and x.

[0052] Preferably, step S4: determining the similarity response map of the spiking feature sequences of the template image and the current frame image includes:

[0053] Obtain the spiking feature sequence of the template image; obtain the spiking feature sequence of the current frame image;

[0054] Determine the temporal similarity and potential similarity between the spiking feature sequence of the template image and the spiking feature sequence of the current frame image, and responsemap As the similarity response map of the pulse feature sequences of the template image and the current frame image.

[0055] An unmanned surface vehicle single-target tracking device based on a spiking neural network provided by the present invention, the device includes:

[0056] End-to-end target tracking model module: configured to construct an end-to-end target tracking model based on a spiking neural network;

[0057] Template pulse feature sequence extraction module: configured to obtain a template image of a target to be tracked and extract the pulse feature sequence of the template image;

[0058] Current frame pulse feature sequence extraction module: configured to obtain the current frame image and extract the pulse feature sequence of the current frame image;

[0059] Similarity determination module: configured to determine the similarity response map of the pulse feature sequences of the template image and the current frame image;

[0060] Tracking result module: configured to determine the target tracking result of the current frame image based on the similarity response map;

[0061] Determination module: configured to read the next frame image from the video stream as the current frame image and trigger the extraction of the current frame pulse feature sequence.

[0062] A computer-readable storage medium provided by the present invention, wherein multiple instructions are stored in the storage medium; the multiple instructions are used to be loaded and executed by a processor to perform the method as described above.

[0063] An electronic device provided by the present invention, characterized in that the electronic device includes:

[0064] A processor for executing multiple instructions;

[0065] A memory for storing multiple instructions;

[0066] Wherein, the multiple instructions are used to be stored by the memory and loaded and executed by the processor to perform the method as described above.

[0067] The present invention can analyze the image sequences collected by the visible light camera and the infrared camera on the unmanned surface vehicle, use the spiking neural network to perform feature extraction and similarity estimation on a target of interest, and realize the function of real-time target tracking in video data.

[0068] The beneficial technical effects brought by the present invention:

[0069] (1) The present invention applies a spiking neural network to the task of single-object tracking on the water surface of an unmanned boat. By virtue of the low-power consumption characteristic during its computational inference process, it solves the difficulty of applying deep neural networks in resource-constrained unmanned boat systems at the algorithm level;

[0070] (2) The present invention designs a conversion method of a spiking neural network for a twin neural network for target tracking, avoiding the difficulty of directly training a spiking neural network;

[0071] (3) The two-stage periodic encoding method and the similarity estimation methods based on direct timing, window timing, and potential proposed by the present invention reduce the inference time of the spiking neural network and support real-time inference. Description of the Drawings

[0072] Figure 1 It is a schematic flow diagram of the method for single-object tracking on the water surface of an unmanned boat based on a spiking neural network according to the present invention;

[0073] Figure 2 It is another schematic flow diagram of the method for single-object tracking on the water surface of an unmanned boat based on a spiking neural network according to the present invention;

[0074] Figure 3 It is a schematic flow diagram of the process of constructing an end-to-end target tracking model based on a spiking neural network according to the present invention;

[0075] Figure 4 It is a schematic diagram of the model of the method for single-object tracking on the water surface of an unmanned boat based on a spiking neural network according to the present invention. Detailed Embodiments

[0076] The present invention will be described in detail below with reference to the drawings and embodiments.

[0077] As Figure 1 - Figure 2 shown, the present invention proposes a method for single-object tracking on the water surface of an unmanned boat based on a spiking neural network, and the method includes:

[0078] Step S1: Construct an end-to-end target tracking model based on a spiking neural network;

[0079] Step S2: Obtain a template image of the target to be tracked, and extract the spiking feature sequence of the template image;

[0080] Step S3: Obtain the current frame image, and extract the spiking feature sequence of the current frame image;

[0081] Step S4: Determine the similarity response map of the spiking feature sequences of the template image and the current frame image;

[0082] Step S5: Based on the similarity response map, determine the target tracking result of the current frame image;

[0083] Step S6: Read the next frame image from the video stream as the current frame image, and enter Step S3; if there is no next frame image, the method ends.

[0084] There are difficulties such as non-differentiability in directly training a spiking neural network, and the implementation accuracy of some unsupervised training methods is difficult to meet the requirements. Therefore, first train a deep neural network, and then use a conversion method to construct an end-to-end object tracking model based on a spiking neural network.

[0085] Furthermore, as Figure 3 - Figure 4 shown, Step S1: Construct an end-to-end object tracking model based on a spiking neural network, including:

[0086] Step S11: Modify the SiamFC model, including: merging all parameters of the batch normalization layer (BatchNormalization, BN) of the SiamFC model into the weight of the convolutional layer in the previous layer; the calculation formula and parameter merging formula of the batch normalization layer are:

[0087]

[0088]

[0089]

[0090] where x is the input vector of the BN layer, BN[x] is the output vector of the BN layer, μ and σ are the mean and variance of the training data respectively, and β and γ are variables obtained by training through gradient descent; is the weight of the l-th convolutional kernel after modification, is the weight coefficient of the l-th BN layer, is the variance of the training data of the l-th layer, is the weight of the l-th convolutional kernel; is the bias of the l-th convolutional layer after modification, is the bias of the l-th convolutional layer, is the mean of the training data of the l-th layer, is the constant coefficient of the l-th BN layer;

[0091] Delete the bias of the last layer of the backbone convolutional neural network of the SiamFC model.

[0092] In this embodiment, all the parameters of the Batch Normalization (BN) layer are incorporated into its preceding convolutional layer, and the BN layer is not independently trained and stored with parameters, which can improve the computational efficiency during the inference process of the spiking neural network. Then, the bias of the last layer of the backbone convolutional neural network of SiamFC is deleted, in order to shorten the spiking simulation time step to improve the model inference speed. Since the subsequent similarity estimation method is based on the convolutional filtering of the spiking feature sequence, the simulation of the pulse signal for the floating-point bias parameter requires a long time step. Therefore, the bias parameter is deleted before training the modified SiamFC model.

[0093] Step S12: Use the first training set to perform iterative training on the modified SiamFC model for several rounds, and then use the second training set including visible light images and infrared light images to train the modified SiamFC model to obtain the trained modified SiamFC model.

[0094] In this embodiment, first use the ILSVRC2015-VID, OTB-2015, and VOT-2018 object tracking datasets to pre-train the modified SiamFC model and stop after 10 training epochs. Then use the visible light and infrared light video data of the Singapore Maritime Dataset (SMD) to alternately train the model, and after another 40 epochs of training, obtain the trained modified SiamFC model, which has the ability to simultaneously analyze visible light and infrared light image sequences.

[0095] Step S13: Convert the trained modified SiamFC model into a spiking neural network model, including:

[0096] Calculate the response value based on the pixel values of each pixel point of the image as the input of the spiking neural network model and the convolutional kernel parameters of the first layer of the backbone convolutional neural network of the trained modified SiamFC model, and input the response value as the fixed potential value of the first layer of the spiking neural network of the spiking neural network model; the calculation formula of the response value is:

[0097]

[0098] Where, is the output response value of the first layer of the spiking neural network, V th is the potential threshold, is the weight of the first layer of the spiking neural network, x j is the input value of the first layer of the spiking neural network, is the bias of the first layer of the spiking neural network;

[0099] For the backbone convolutional neural network of the trained modified SiamFC model, from the second layer to the last layer, the parameter normalization method is used layer by layer to convert the weights of the backbone convolutional neural network into the weights of the spiking convolutional network of the spiking neural network model. The conversion formula is:

[0100]

[0101] where the normalization factor λ l is the 99.9% largest real number in the activation value vector of the l-th layer, w l , b l are the weights and biases of the convolutional layer of the l-th layer respectively, are the weights and biases of the converted spiking convolutional network respectively;

[0102] The backbone convolutional neural network after SiamFC conversion is the spiking convolutional network, and its output features are used as the spiking feature sequence of the spiking neural network model.

[0103] In this embodiment, first, the input layer encoding is implemented. Since the time step required for frequency encoding is relatively long, which will reduce the model inference speed, the response value is calculated by multiplying the pixel value of the input image with the convolutional kernel parameters of the first layer of the backbone convolutional neural network of SiamFC, and used as the fixed potential value input of the first layer of the spiking neural network. Then, the spiking neuron is implemented using the integrate-and-fire model and the soft reset method, that is, after outputting a spiking signal, the neuron potential only decreases by 1 to prevent the neuron from losing too much information. Finally, the parameter normalization method is used layer by layer to convert the weights of the backbone convolutional neural network into the weights of the spiking convolutional network. The normalization factor is the 99.9% largest real number in the activation value vector of each layer, avoiding the problem of too high firing rate of shallow spiking neurons and too low firing rate of deep spiking neurons.

[0104] Step S14: Generate the potential threshold by means of two-stage periodic encoding, and use the potential threshold as the potential threshold of the spiking convolutional network, including:

[0105] The two-stage periodic encoding refers to encoding the spiking feature sequence in a two-stage threshold manner, which is divided into two potential thresholds; based on the two potential thresholds, the potential threshold of the spiking neuron is periodically changed. The formula for calculating the potential threshold is:

[0106]

[0107] where p is the threshold change period, t is the current spiking simulation time, and ε is the activation threshold.

[0108] In this embodiment, two potential thresholds are used to periodically change the threshold of the spiking neuron, replacing the original frequency encoding of the spiking neural network to improve the model performance. If the spike trains are too randomly distributed and chaotic, a large number of spike signals cannot be fully utilized, thus prolonging the time steps required for model inference.

[0109] Step S15: Establish a similarity calculation formula for the spiking neural network model, where:

[0110] Similarity refers to the similarity of the spike feature sequences between the target template and the search region of the current frame, including temporal similarity and potential similarity;

[0111] The temporal similarity considers the direct temporal correlation and window temporal correlation of the spike feature sequences. The direct temporal correlation directly calculates the convolution result of the spike feature sequences at each time step. Considering that the correlation of spike signals is lower for larger time intervals, the window temporal correlation determines the temporal similarity according to the convolution result of the signal values of two spike sequences within a sliding time window (2τ); the formula for calculating the temporal similarity is as follows:

[0112]

[0113] where z and x represent the target template image and the search region image of the current frame, respectively, represents the spike convolution network, τ is the time window parameter, T is the total simulation steps of the spiking neural network, and k is the weight set by the experiment;

[0114] The potential similarity considers the overall potential correlation of the spike feature sequences and calculates the convolution result of the comprehensive spike signals in all simulation time steps. The calculation formula is as follows:

[0115]

[0116] where, is the bias.

[0117] Step S16: Generate a response map of the spiking neural network model, including:

[0118] Based on the temporal similarity and potential similarity between the target template and the search region of the current frame, linearly weighted fusion of the temporal similarity and potential similarity is performed to obtain a two-dimensional matrix of spike feature sequence similarity. The calculation formula is

[0119] f s (z, x) = f P (z, x) + λf T (z, x),

[0120] where λ is the weight set by the experiment;

[0121] Then, traverse the similarity two-dimensional matrices generated in each batch of the inference results of the spiking neural network model this time, and select the matrix with the largest similarity value; subtract the minimum element value in the matrix from each element in the matrix, and divide by the sum of all elements in the matrix for normalization; finally, linearly fuse the hanning window matrix with the similarity matrix to penalize the points at the coordinate positions far from the target, and use the obtained fusion result as the response map of the image input to the spiking neural network model. The calculation formula is as follows:

[0122] response map =(1 - α)·f(z, x)+α·f hanning

[0123] where f is the normalized similarity matrix, f hanning is the hanning window matrix, α is the weight set by the experiment, and f(z, x) is the similarity matrix of z and x after normalization.

[0124] In this embodiment, first, linearly weighted fusion is performed on the above two spiking feature correlations to obtain the final similarity estimation result, which is beneficial for the model to achieve the ideal inference accuracy within a shorter time step. The calculation formula is as follows:

[0125] f s (z, x)=f P (z, x)+λf T (z, x),

[0126] Then, traverse the two-dimensional similarity matrices generated in each batch of the inference results this time, and select the matrix with the point having the largest similarity; subtract the minimum element value in the matrix from each element in the matrix, and divide by the sum of all elements in the matrix for normalization. Finally, linearly fuse the hanning window matrix with the similarity matrix to penalize the points at the coordinate positions far from the target, and obtain the response map. The calculation formula is as follows:

[0127] response map =(1 - α)·f(z, x)+α·f hanning

[0128] Further, the step S4: determining the similarity response map of the pulse feature sequences of the template image and the current frame image includes:

[0129] Obtain the pulse feature sequence of the template image; obtain the pulse feature sequence of the current frame image;

[0130] Determine the temporal similarity and potential similarity between the pulse feature sequence of the template image and the pulse feature sequence of the current frame image, and set response map as the similarity response map of the pulse feature sequences of the template image and the current frame image.

[0131] The present invention provides a specific embodiment of an unmanned surface vehicle single-object tracking method based on a spiking neural network.

[0132] S1: Construct an end-to-end object tracking model based on a spiking neural network;

[0133] There are difficulties such as non-differentiability in directly training a spiking neural network, and the implementation accuracy of some unsupervised training methods is difficult to meet the requirements. Therefore, first train a deep neural network, and then use a conversion method to construct an end-to-end object tracking model based on a spiking neural network, as Figure 2 shown. The specific steps are as follows:

[0134] S1_1: Modify the SiamFC model.

[0135] First, incorporate all the parameters of the batch normalization layer (Batch Normalization, BN) into its preceding convolutional layer, and do not independently train and store the parameters of the BN layer to improve the computational efficiency during the inference process of the spiking neural network. The calculation formula of the BN layer and the parameter merging formula are as follows:

[0136]

[0137]

[0138]

[0139] where μ and σ are the mean and variance of the training data respectively, and β and γ are variables obtained by training through gradient descent.

[0140] Then, delete the bias of the last layer of the backbone convolutional neural network of SiamFC, that is, the Conv5 convolutional layer only has convolutional kernel parameters. This is to shorten the pulse simulation time step to improve the model inference speed. Since the subsequent similarity estimation method is based on the convolutional filtering of the pulse feature sequence, the simulation of the pulse signal for the floating-point bias parameter requires a long time step. For example, representing 0.007 requires outputting 7 pulse signals within 1000 time steps, which greatly increases the model inference speed. Therefore, the bias parameter is deleted before training. The structure of the backbone convolutional neural network is as follows:

[0141] Convolutional layer Convolution kernel size Stride Number of channels Conv1 11×11 2 96 maxpool1 3×3 2 96 Conv2 5×5 1 256 maxpool2 3×3 2 256 Conv3 3×3 1 192 Conv4 3×3 1 192 Conv5 3×3 1 128

[0142] S1_2: Train the model using the dataset of object detection and tracking for surface unmanned vessels.

[0143] First, pre-train the model using the object tracking datasets of ILSVRC2015-VID, OTB-2015, and VOT-2018. The initial learning rate is 0.01, the learning rate decay method is exponential, and the decay coefficient is 0.87. The optimizer uses momentum gradient descent with a coefficient of 0.9. During training, the resolution of the target template image and the search area image of the current frame are 127×127 and 255×255 respectively, and the batch_size is 8. The graphics card used for training is NVIDIA RTX2080, and the training stops after 10 training epochs.

[0144] Then, alternately train the model using the visible light and infrared light video data of the Singapore Maritime Dataset (SMD). The visible light videos include data taken from the shore and on board, with a total of 51 videos; the infrared videos are taken from the shore, with a total of 30 videos. The training configuration is the same as that during pre-training. After training for another 40 epochs, a target tracking model is obtained, which has the ability to analyze visible light and infrared light image sequences simultaneously.

[0145] S1_3: Convert the trained SiamFC model into a spiking neural network;

[0146] First, implement the input layer encoding. Since the time steps required for frequency encoding are relatively long, which will reduce the model inference speed, the response value is calculated by multiplying the pixel values of the input image with the convolution kernel parameters of the first layer of the backbone convolutional neural network of SiamFC, and used as the fixed potential value input of the first layer of the spiking neural network. The calculation formula is as follows:

[0147]

[0148] Then, use the integrate-and-fire model and the soft reset method to implement the spiking neuron, that is, after outputting a pulse signal, the neuron potential only decreases by 1 to prevent the neuron from losing too much information. The neuron potential accumulation method and the output pulse calculation formula are as follows:

[0149]

[0150]

[0151]

[0152] Where represents the potential accumulated by the spiking neuron i of the l-th layer at time t, represents the output pulse binary signal, with a value of 0 or 1, and U(·) represents the unit step function, whose function value is 1 when the input is greater than 0, otherwise 0.

[0153] Finally, the parameter normalization method is used layer by layer to convert the weights of the backbone convolutional neural network into the weights of the spiking convolutional network. The calculation formula is as follows:

[0154]

[0155] Normalization factor λ l is the 99.9% largest real number in the activation value vector of each layer, avoiding the problems of too high firing rates of shallow spiking neurons and too low firing rates of deep spiking neurons.

[0156] S1_4: Implement a two-stage periodic coding method to replace the frequency coding of the spiking neural network.

[0157] By periodically changing the potential threshold of two types of spiking neurons to replace the original frequency coding of the spiking neural network, the model performance can be improved. Since the subsequent similarity estimation method needs to calculate the temporal similarity of the spiking feature sequence using convolutional operations, if the spiking sequence is too randomly distributed and chaotic, a large number of spiking signals cannot be fully utilized, resulting in an extended time step required for model inference. Therefore, it is necessary to make the generated spiking feature sequence arranged in an orderly manner so that each spiking signal can be fully utilized. The specific calculation formula is as follows:

[0158]

[0159] According to experimental optimization, the model performance is optimal when the period p = 5 and the activation threshold ε = 0.6. That is, the potential threshold of the spiking neuron alternates every 5 time steps between 0.6 and positive infinity.

[0160] S1_5: Implement a similarity estimation method based on time series and potential;

[0161] SiamFC uses convolutional operations to calculate the similarity matrix between the target template feature and the feature of the current frame search area. Since it is single-object tracking, it is considered that the point with the largest element value in the similarity matrix is the position of the target in the current frame. The calculation formula is as follows:

[0162]

[0163] where z and x represent the target template image and the current frame search area image respectively, denotes that the backbone convolutional neural network is used to extract features. For the converted spiking neural network model, similarity estimation methods based on time series and potential are designed respectively to calculate the similarity matrix.

[0164] The above-mentioned temporal similarity considers the direct temporal correlation and window temporal correlation of the pulse feature sequence. The direct temporal correlation directly calculates the convolution result of the pulse feature sequence at each time step. The window temporal correlation takes into account that the correlation of pulse signals with a greater time interval is lower, and determines the temporal similarity based on the convolution result of the signal values of two pulse sequences within a sliding time window (2τ). The formula for calculating the temporal similarity is as follows:

[0165]

[0166] According to experimental optimization, when the time window parameter τ = 1 (i.e., the time window is 2 time steps), the model performance is optimal.

[0167] The similarity estimation based on potential considers the overall potential correlation of the pulse feature sequence, and calculates the convolution result of the integrated pulse signals in all simulated time steps. The formula is as follows:

[0168]

[0169] S1_6: Implement the output layer to decode and generate a response map to obtain an end-to-end object tracking model.

[0170] First, linearly weighted fusion is performed on the above two pulse feature correlations to obtain the final similarity estimation result, which is beneficial for the model to achieve ideal inference accuracy within a shorter time step. The formula is as follows:

[0171] f s (z, x) = f P (z, x) + λf T (z, x),

[0172] According to experimental optimization, when the weighted parameter λ = 5.4, the model performance is optimal.

[0173] Then, traverse the two-dimensional similarity matrix generated in each batch of this inference result, and select the matrix with the point having the maximum similarity; subtract the minimum element value in the matrix from each element in the matrix, and divide by the sum of all elements in the matrix for normalization. The formula is as follows:

[0174]

[0175] Finally, linear fusion is performed between the hanning window matrix and the similarity matrix to penalize the points at coordinate positions far from the target, and a response map is obtained. The formula is as follows:

[0176] response map = (1 - α)·f(z, x) + α·f hanning

[0177] According to the experimental optimization, the model performance is optimal when the linear parameter α = 1.8. After the construction of the end-to-end object tracking model based on the spiking neural network is completed, it is used for the analysis of visible light and infrared video data collected in subsequent actual tasks.

[0178] S2: Obtain the target template image to be tracked and extract its spiking feature sequence.

[0179] In the actual application task, the target template image to be tracked can be automatically obtained by the target detection algorithm or manually marked. First, obtain the target bounding box (x, y, w, h) and the entire image. After expanding the pixel interval length of p around the target bounding box and then intercepting, scale the intercepted area to a resolution of 127×127 to obtain the target template image. The calculation formula is as follows:

[0180] s(w + 2p)·s(h + 2p) = A

[0181] where s is the scaling ratio, A = 127 2 is the resolution size of the target template image, and p = (w + h) / 4 is the interval length of the expanded area around the target bounding box. Then, input the target template image into the spiking convolutional network to obtain the spiking feature sequence and cache it in the video memory, as shown in the upper branch of Figure 3 . When calculating subsequent image frames, the target template image will not perform repeated feature extraction operations until the target template image is updated. The target template image supports visible light images and infrared images, corresponding to day and night scenes.

[0182] S3: Obtain the current frame image, scale and intercept it, and then extract its spiking feature sequence.

[0183] Obtain the current frame image from the sensor video data. With the position of the previous frame of the target as the center, intercept the current frame search area with a resolution of 255×255 according to the scaling ratio s obtained in the previous step. Then, obtain the current frame search area image tensor with batch_size = 3 according to three preset prior target sizes, and input it into the spiking convolutional network to obtain the spiking feature sequence, as shown in the lower branch of Figure 3 .

[0184] S4: Calculate the similarity response map of the spiking feature sequences of the target template and the current frame search area;

[0185] Read the spiking feature sequence of the target template in the video memory, and calculate the similarity response map response of its spiking feature sequence with the current frame search area according to the mathematical model implemented in steps S1_5 and S1_6 map .

[0186] S5: Calculate the target tracking result of the current frame;

[0187] Obtain the relative coordinates of the maximum response value in the response map, calculate the block distance between it and the center of the response map, and suppress the maximum distance from the center point, which should not exceed half of the width value of the response map. Decode the relative coordinates to obtain the relative coordinate offset in the current frame search area of size 255×255. Then, calculate the target position in the current frame based on the target position (x, y, w, h) of the previous frame.

[0188] S6: Read the next frame of the image and jump to step S3; if there is no subsequent input image, end the target tracking process.

[0189] The present invention provides an unmanned surface vehicle single-target tracking device based on a spiking neural network, including:

[0190] End-to-end target tracking model module: Configured to construct an end-to-end target tracking model based on a spiking neural network;

[0191] Template spiking feature sequence extraction module: Configured to obtain a template image of the target to be tracked and extract the spiking feature sequence of the template image;

[0192] Current frame spiking feature sequence extraction module: Configured to obtain the current frame image and extract the spiking feature sequence of the current frame image;

[0193] Similarity determination module: Configured to determine the similarity response map of the spiking feature sequences of the template image and the current frame image;

[0194] Tracking result module: Configured to determine the target tracking result of the current frame image based on the similarity response map;

[0195] Determination module: Configured to read the next frame of the image from the video stream as the current frame image and trigger the extraction of the current frame spiking feature sequence.

[0196] A computer-readable storage medium provided by the present invention, in which multiple instructions are stored; the multiple instructions are used to be loaded and executed by a processor to perform the method as described above.

[0197] An electronic device provided by the present invention, characterized in that the electronic device includes:

[0198] A processor for executing multiple instructions;

[0199] A memory for storing multiple instructions;

[0200] Wherein, the multiple instructions are used to be stored by the memory and loaded and executed by the processor to perform the method as described above.

[0201] The above specific embodiments only describe the design principle of the present invention. The shapes and names of the components in this description can be different and are not restricted. Therefore, those skilled in the art of the present invention can modify or make equivalent substitutions for the technical solutions recorded in the foregoing embodiments; and these modifications and substitutions do not depart from the gist and technical solutions of the present invention, and shall all fall within the protection scope of the present invention.

Claims

1. A method for single target tracking on the water surface of an unmanned boat based on a spiking neural network, characterized in that, The method includes: Step S1: Construct an end-to-end target tracking model based on a spiking neural network; Step S2: Obtain a template image of the target to be tracked, and extract the spiking feature sequence of the template image; Step S3: Obtain the current frame image, and extract the spiking feature sequence of the current frame image; Step S4: Determine the similarity response map of the spiking feature sequences of the template image and the current frame image; Step S5: Based on the similarity response map, determine the target tracking result of the current frame image; Step S6: Read the next frame image from the video stream as the current frame image, and enter Step S3; if there is no next frame image, the method ends.

2. The method according to claim 1, characterized in that, The Step S1: Construct an end-to-end target tracking model based on a spiking neural network, includes: Step S11: Modify the SiamFC model, including: merging all parameters of the batch normalization layer (BatchNormalization, BN) of the SiamFC model into the convolutional layer weights of the previous layer; the calculation formula and parameter merging formula of the batch normalization layer are: Among them, x is the input vector of the BN layer, BN[x] is the output vector of the BN layer, μ and σ are the mean and variance of the training data respectively, and β and γ are variables obtained through gradient descent training; is the weight of the l-th convolutional kernel after modification, is the weight coefficient of the l-th BN layer, is the variance of the training data of the l-th layer, is the weight of the l-th convolutional kernel; is the bias of the l-th convolutional layer after modification, is the bias of the l-th convolutional layer, is the mean of the training data of the l-th layer, is the constant coefficient of the l-th BN layer; Delete the bias of the last layer of the backbone convolutional neural network of the SiamFC model; Step S12: Use the first training set to perform iterative training on the modified SiamFC model for several rounds, and then use the second training set including visible light images and infrared light images to train the modified SiamFC model to obtain the trained modified SiamFC model; Step S13: Convert the trained modified SiamFC model into a spiking neural network model; Step S14: Generate a potential threshold through a two-stage periodic encoding method, and use the potential threshold as the potential threshold of the spiking convolutional network, including: The two-stage periodic encoding refers to encoding the spiking feature sequence in a two-stage threshold manner, which is divided into two potential thresholds; based on the two potential thresholds, periodically change the potential threshold of the spiking neuron, and the formula for calculating the potential threshold is: where p is the threshold change period, t is the current spiking simulation time, and ε is the activation threshold; Step S15: Establish the similarity calculation formula of the spiking neural network model, where: Similarity refers to the similarity of the spiking feature sequences between the target template and the current frame search area, including temporal similarity and potential similarity; The temporal similarity considers the direct temporal correlation and window temporal correlation of the spiking feature sequence. The direct temporal correlation directly calculates the convolution result of the spiking feature sequence at each time step, and the window temporal correlation takes into account that the correlation of spiking signals with a farther time interval is lower, and determines the temporal similarity according to the convolution result of the signal values of two spiking sequences within the sliding time window 2τ; the calculation formula of the temporal similarity is as follows: where z and x represent the target template image and the current frame search region image respectively, denotes the pulsed convolutional network, τ is the time window parameter, T is the total simulation steps of the spiking neural network, and k is the weight set by the experiment; The potential similarity considers the overall potential correlation of the spiking feature sequence, and calculates the convolution result of the comprehensive spiking signals in all simulation time steps, and the calculation formula is as follows: Among them, is the bias. Step S16: Generate the response map of the spiking neural network model.

3. The method according to claim 2, characterized in that The Step S13: Convert the trained modified SiamFC model into a spiking neural network model, includes: Calculating a response value based on the pixel values of each pixel point of the image that is the input of the spiking neural network model and the convolutional kernel parameters of the first layer of the backbone convolutional neural network of the trained modified SiamFC model, and using the response value as the input of the fixed potential value of the first layer of spiking neural network of the spiking neural network model; the calculation formula of the response value is: Among them, is the output response value of the first-layer spiking neural network, V th is the potential threshold, are the weights of the first-layer spiking neural network, x j is the input value of the first-layer spiking neural network, is the bias of the first-layer spiking neural network; For the backbone convolutional neural network of the trained modified SiamFC model, from the second layer to the last layer, the parameter normalization method is used layer by layer to convert the backbone convolutional neural network weights into the weights of the spiking convolutional network of the spiking neural network model, and the conversion formula is: Among them, the normalization factor λ l is the 99.9% largest real number in the activation value vector of the l-th layer, and w l , b l are the weights and biases of the convolutional layer of the l-th layer respectively, are the weights and biases of the converted spiking convolutional network respectively; The backbone convolutional neural network after SiamFC conversion is the spiking convolutional network, and its output features are used as the spiking feature sequence of the spiking neural network model.

4. The method according to claim 2, characterized in that, Step S16: Generating a response map of the spiking neural network model, including: Based on the temporal similarity and potential similarity between the target template and the current frame search area, linearly weighted fusion of the temporal similarity and potential similarity is performed to obtain a two-dimensional matrix of pulse feature sequence similarity. The calculation formula is f s (z,x) = f P (z,x) + λf T (z,x), where λ is the weight set by the experiment; Then traverse the similarity two-dimensional matrix generated by each batch in the inference result of the spiking neural network model this time, and select the matrix with the largest similarity value; subtract the minimum element value in the matrix from each element in the matrix, and divide by the sum of all elements in the matrix for normalization operation; finally, linearly fuse the hanning window matrix and the similarity matrix to punish the points at the coordinate positions far from the target, and use the fusion result as the response map of the image input to the spiking neural network model. The calculation formula is as follows: response map = (1 - α)·f(z, x) + α·f hanning Among them, f is the normalized similarity matrix, and f hanning is the Hanning window matrix, α is the weight set by the experiment, and f(z, x) is the similarity matrix of z and x after normalization.

5. The method according to any one of claims 1 to 4, characterized in that, The step S4: Determining the similarity response map of the pulse feature sequences of the template image and the current frame image, including: Obtaining the pulse feature sequence of the template image; obtaining the pulse feature sequence of the current frame image; Determine the timing similarity and potential similarity between the pulse feature sequence of the template image and the pulse feature sequence of the current frame image, and use response map as the similarity response map of the pulse feature sequences of the template image and the current frame image.

6. An unmanned surface vehicle single-target tracking device based on a spiking neural network, characterized in that, The device includes: An end-to-end target tracking model module: configured to construct an end-to-end target tracking model based on a spiking neural network; A template pulse feature sequence extraction module: configured to obtain a template image of a target to be tracked and extract the pulse feature sequence of the template image; A current frame pulse feature sequence extraction module: configured to obtain a current frame image and extract the pulse feature sequence of the current frame image; A similarity determination module: configured to determine the similarity response map of the pulse feature sequences of the template image and the current frame image; A tracking result module: configured to determine the target tracking result of the current frame image based on the similarity response map; A determination module: configured to read the next frame image from the video stream as the current frame image and trigger the extraction of the current frame pulse feature sequence.

7. A computer-readable storage medium, in which multiple instructions are stored; the multiple instructions are used to be loaded and executed by a processor to perform the method according to any one of claims 1-5.

8. An electronic device, characterized in that The electronic device includes: A processor for executing multiple instructions; A memory for storing multiple instructions; wherein the multiple instructions are used to be stored by the memory and loaded and executed by the processor to perform the method according to any one of claims 1-5.