Hash retrieval method based on spiking neural network and soft similarity loss
By introducing Transformer feature extraction layer and hash layer into the pulsed neural network, combined with soft similarity loss, the problem of high computing cost and energy consumption of pulsed neural networks in DVS data retrieval is solved, and efficient retrieval and performance improvement of multiple types of data is achieved.
Patent Information
- Application Number
- CN202510040306.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-10
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2045-01-10
AI Technical Summary
Current pulsed neural networks have high computational cost and energy consumption when processing DVS data, and lack effective hashing methods to achieve efficient multi-class data retrieval.
Using a hash retrieval method based on pulsed neural network and soft similarity loss, the hash retrieval method is used to construct the Transformer feature extraction layer and hash layer, combined with the soft similarity matrix, the hash code generation is optimized, and the calculation amount and energy consumption are reduced.
It realizes efficient retrieval of multiple categories of data, improves retrieval performance, reduces calculation amount and energy consumption, and can more accurately reflect the similarity differences between categories.
Smart Images

Figure BDA0005236733940000081 
Figure BDA0005236733940000082 
Figure BDA0005236733940000091
Abstract
Description
Technical Field
[0001] The present invention relates to a hash retrieval method, in particular to a hash retrieval method based on a spiking neural network and soft similarity loss. Background Art
[0002] In recent years, with the development of dynamic vision sensor (DVS) technology, DVS data has achieved microsecond-level time resolution and efficient capture of fast dynamic scenes by converting the brightness change of each pixel into an event stream, becoming an important research object in the field of computer vision. However, with the rapid growth of DVS data, retrieval networks based on traditional artificial neural networks face severe challenges in computing cost and energy consumption when processing this massive data, restricting their practical applications.
[0003] As the third generation of neural networks, spiking neural networks transmit information by using binary spikes, showing lower energy consumption and higher computing efficiency, and becoming an effective solution for processing DVS data. However, there are some key problems when current spiking neural networks are applied to retrieval tasks: (1) When spike neurons in a spiking neural network convert floating-point membrane potentials into binary spikes, information loss will occur, thus restricting the performance of the spiking neural network; therefore, although the membrane potential contains richer feature information, previous spiking neural network methods have not fully utilized it; (2) Deep hashing is an important technology for large-scale data retrieval, with high efficiency in retrieval speed and storage. Although large-scale DVS data urgently requires an efficient retrieval network, no hashing method has been proposed in spiking neural networks; traditional deep hashing methods usually generate a binary similarity matrix as the target through class labels to construct similarity loss. If samples belong to the same class, the value in the binary similarity matrix is 1, otherwise 0. However, this predefined binary similarity matrix ignores the potential similarity differences between classes.
[0004] The above key problems limit the performance of current spiking neural networks in DVS data retrieval tasks, and new solutions are urgently needed to improve their efficiency and performance. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a hash retrieval method based on a spiking neural network and soft similarity loss, which can reduce the computational amount and energy consumption of the retrieval network, achieve efficient retrieval of multiple data forms, and has relatively high retrieval performance of the hash retrieval model.
[0006] The technical solution adopted by the present invention to solve the above technical problem is: a hash retrieval method based on a spiking neural network and soft similarity loss, comprising the following steps:
[0007] Step 1: Select N data from the original data set as training samples, and perform data preprocessing on the training samples to obtain preprocessed training data;
[0008] Step 2: Construct a hash retrieval model to be trained, including a Transformer feature extraction layer based on a spiking neural network, a hash layer based on a spiking neural network, and a classification head. Randomly shuffle the N preprocessed training data and input them into the Transformer feature extraction layer based on a spiking neural network to obtain the sample features after feature extraction.
[0009] Step 3: Input the sample features after feature extraction into the hash layer based on the pulse neural network to obtain the membrane potential features of the training samples and the hash codes corresponding to the training samples. Then, input the hash codes corresponding to the training samples into the classification head, and the classification head outputs the classification results of the training samples.
[0010] Step 4: Calculate the cosine similarity between the membrane potential features of different training samples according to the membrane potential features of different training samples, accumulate and normalize the cosine similarities between the membrane potential features of different training samples, obtain the category similarity between different categories of training samples, and then form a soft similarity matrix with the category similarities between different categories of training samples, define the loss function of the hash retrieval model to be trained, obtain the value of the loss function through the soft similarity matrix, the category label of the training sample, the hash code corresponding to the training sample and the classification result of the training sample, start the training process by iteratively updating the hash retrieval model to be trained through back propagation, close the hash codes generated by similar training sample pairs, and distance the hash codes generated by dissimilar training sample pairs, until the training process ends after the set maximum number of iterative updates is reached, and the trained hash retrieval model is obtained;
[0011] Step 5: Select M data from the target data set as query samples and form a query set, use the remaining data in the target data set as retrieval samples and form a retrieval set, preprocess the query samples to obtain preprocessed query samples, input the preprocessed query samples into the trained hash retrieval model to obtain the hash code corresponding to the query samples; preprocess the retrieval samples to obtain preprocessed retrieval samples, input the preprocessed retrieval samples into the trained hash retrieval model to obtain the hash code corresponding to the retrieval samples, and define the set consisting of the hash codes corresponding to all retrieval samples as a hash code retrieval library;
[0012] Step 6: Search the hash code retrieval library for the hash code with the closest Hamming distance to the hash code corresponding to the query sample, and use the retrieval sample corresponding to the hash code as the retrieval result to complete the retrieval process for the query sample.
[0013] Compared with the prior art, the advantages of the present invention are as follows: First, by utilizing the binary pulse characteristics of the spiking neural network, a Transformer feature extraction layer based on the spiking neural network and a hash layer based on the spiking neural network are constructed. It is applicable not only to static images but also can efficiently process DVS data, reduce the computational complexity and energy consumption of the retrieval network, and achieve efficient retrieval of multiple data forms; after extracting sample features through the Transformer feature extraction layer based on the spiking neural network, different from the traditional deep hashing method that uses a sign function to construct a quantization layer in the past, the present invention uses spiking neurons to construct a hash layer. Combining the pulse characteristics and the binary characteristics of the hash code, without an additional quantization module, the final hash code is obtained through decision-making at multiple time steps; when optimizing the hash retrieval network, a dynamic soft similarity loss is introduced. A learnable soft similarity matrix is constructed through the membrane potential characteristics in the spiking neurons, which is continuously iteratively updated during the training process to capture more complex relationships between categories, enabling effective information to be retained in the feature learning of different training rounds; compared with the traditional hard similarity loss, this dynamic soft similarity loss can more accurately reflect the similarity differences between categories, compensate for the information loss caused by the spiking neural network, and improve the retrieval performance of the model by jointly optimizing the dynamic soft similarity loss and the classification loss.
[0014] Specifically, in step 1, when the original data set is a dynamic vision sensor (DVS) data set, the specific process of preprocessing the training samples is as follows: Integrate the files of the dynamic vision sensor (DVS) data set according to the time window to obtain 8 video data, then the preprocessed training data includes video data of 8 time steps.
[0015] When the original data set is a static image data set, the specific process of preprocessing the training samples is as follows: Copy a single image in the static image data set 4 times and arrange the copied images in order to form a 4-time-step image sequence, which is used as the preprocessed training data.
[0016] Specifically, the specific process of step 3 is as follows:
[0017] Step 3-1: The hash layer based on the spiking neural network includes a pooling layer, a linear layer, spiking neurons, and a hash conversion layer. Input the sample features after feature extraction into the pooling layer, and the pooling layer applies global average pooling to obtain a pooled feature vector f with a dimension of 8×512×1.
[0018] Step 3-2: Input f into the linear layer, and through the mapping of the linear layer, obtain a floating-point value feature x with the same length L as the preset hash code. The dimension of x is 8×L, where 8 is the time step length. Then input x into the spiking neurons to obtain binary features b at multiple time steps and the membrane potential feature m of the training sample.
[0019] Step 3-3: Input b into the hash conversion layer. The hash conversion layer performs an AND operation on b in multiple time step dimensions to obtain the binary hash code h, where h = And(SN((W T f + o))), where W represents the weight matrix of the linear layer, W T represents the transpose of W, o represents the bias of the linear layer, W T f + o = x, SN(…) represents the spiking neuron, and And(…) represents the AND operation performed by the hash conversion layer;
[0020] Step 3-4: Input h into the classification head, and the classification head outputs the classification result p.
[0021] Specifically, the specific process of step 4 is as follows:
[0022] Step 4-1: Define the loss function L of the hash retrieval model to be trained, where L = α·L s + β·L cls , where L cls is the classification loss, C is the total number of categories of the training samples input in the current training round, t represents the category number, 1 ≤ t ≤ C, y t represents the category label of the training sample. If the training sample belongs to the t-th category, y t = 1, otherwise y t = 0, and p t represents the probability that the training sample belongs to the t-th category in the classification result output by the classification head;
[0023] L s is the soft similarity loss, where e is the natural constant, h p represents the hash code corresponding to the p-th training sample, h q represents the hash code corresponding to the q-th training sample, is the transpose of h p ; in the first training round, is obtained from the category label of the training sample, t p represents the category number of the p-th training sample, t q represents the category number of the q-th training sample. If the p-th training sample and the q-th training sample belong to the same category, then If they belong to different categories, then Obtain the soft similarity matrix in the first training round, and denote the soft similarity matrix as S soft , S soft is in the initial state of a all-zero matrix at the beginning of training, where s ij ′ is located in S softAn element in the i-th row and j-th column in , indicating the category similarity between the i-th category and the j-th category, s ij ′ The acquisition process of is as follows: calculate the cosine similarity between the membrane potential features of each training sample belonging to the i-th category in the current training round and the membrane potential features of each training sample belonging to the j-th category, then accumulate the cosine similarities between the membrane potential features of all two training samples belonging to the i-th category and the j-th category, and then normalize the result of the cosine similarity accumulation according to the number of training samples in the current training round to obtain s ij ′ , and at the end of the current training round, Update to s obtained in the current training round ij ′ , and used as the next training round in the current training round α and β are used to balance L s and L cls Parameters;
[0024] Step 4-2: Set the maximum number of training rounds, and use the AdamW optimization algorithm to iteratively update the hash retrieval model to be trained according to the loss function K. When the maximum number of training rounds is reached, stop the iterative update process and obtain the trained hash retrieval model.
[0025] Through the constraint of classification loss, the trained hash retrieval model can effectively optimize the inter-class distance of hash codes and generate more discriminative hash codes for samples of different categories; through the constraint of soft similarity loss, the trained hash retrieval model can make full use of the similarity differences between different categories, bring the hash codes generated by similar sample pairs closer, and move the hash codes of dissimilar sample pairs further apart.
[0026] Specifically, in step 4-1, α=0.001, β=1.
[0027] Specifically, in step 4-2, when the original data set adopts the CIFAR10 data set, the maximum number of training rounds is set to 400 times; when the original data set adopts the UCF101-DVS data set or the HMDB51-DVS data set, the maximum number of training rounds is set to 200 times. DETAILED DESCRIPTION
[0028] The present invention is described in further detail below.
[0029] A hash retrieval method based on a pulse neural network and soft similarity loss comprises the following steps:
[0030] Step 1: Select N data from the original dataset as training samples, and perform data preprocessing on the training samples to obtain preprocessed training data. Among them, when the original dataset is a Dynamic Vision Sensor (DVS) dataset, the specific process of preprocessing the training samples is as follows: Integrate the files of the DVS dataset according to the time window to obtain 8 video data, and the preprocessed training data includes video data with 8 time steps;
[0031] When the original dataset is a static image dataset, the specific process of preprocessing the training samples is as follows: Copy a single image in the static image dataset 4 times and arrange the copied images in order to form a sequence of 4 time-step images, which is used as the preprocessed training data.
[0032] Step 2: Construct a hash retrieval model to be trained, including a Transformer feature extraction layer based on a spiking neural network, a hash layer based on a spiking neural network, and a classification head. Randomly shuffle the N preprocessed training data and input them into the Transformer feature extraction layer based on a spiking neural network to obtain sample features after feature extraction.
[0033] Step 3: Input the sample features after feature extraction into the hash layer based on a spiking neural network to obtain the membrane potential features of the training samples and the hash codes corresponding to the training samples. Then input the hash codes corresponding to the training samples into the classification head, and the classification head outputs the classification results of the training samples. The specific process is as follows:
[0034] Step 3-1: The hash layer based on a spiking neural network includes a pooling layer, a linear layer, spiking neurons, and a hash conversion layer. Input the sample features after feature extraction into the pooling layer, and the pooling layer applies global average pooling to obtain a pooled feature vector f with a dimension of 8×512×1.
[0035] Step 3-2: Input f into the linear layer, and through the mapping of the linear layer, obtain a floating-point value feature x with the same length L as the preset hash code. The dimension of x is 8×L, where 8 is the time-step length. Then input x into the spiking neurons to obtain binary features b with multiple time steps and the membrane potential features m of the training samples.
[0036] Step 3-3: Input b into the hash conversion layer, and the hash conversion layer performs an AND operation on b in multiple time-step dimensions to obtain a binary hash code h, h = And(SN((W T f + o))), where W represents the weight matrix of the linear layer, W T represents the transpose of W, o represents the bias of the linear layer, W T f + o = x, SN(…) represents the spiking neurons, and And(…) represents the AND operation performed by the hash conversion layer.
[0037] Step 3-4: Input h into the classification head, and the classification head outputs the classification result p.
[0038] After extracting the sample features through the Transformer feature extraction layer based on the spiking neural network, a hash layer is constructed using spiking neurons. Combining the spiking features and the binary characteristics of the hash code, decisions are made at multiple time steps to obtain the final hash code.
[0039] Step 4: Calculate the cosine similarity between the membrane potential features of different training samples according to the membrane potential features of different training samples, accumulate and normalize the cosine similarity between the membrane potential features of different training samples to obtain the class similarity between different classes of training samples, then form the class similarity between different classes of training samples into a soft similarity matrix, define the loss function of the hash retrieval model to be trained, obtain the value of the loss function through the soft similarity matrix, the class labels of the training samples, the hash codes corresponding to the training samples, and the classification results of the training samples, and start the training process by iteratively updating the hash retrieval model to be trained through backpropagation, pulling closer the hash codes generated by similar training sample pairs and pulling farther apart the hash codes generated by dissimilar training sample pairs until the set maximum number of iterative updates is reached, and then end the training process to obtain the trained hash retrieval model; the specific process is as follows:
[0040] Step 4-1: Define the loss function L of the hash retrieval model to be trained, L = α·L s +β·L cls , where L cls is the classification loss, C is the total number of classes of the training samples input in the current training round, t represents the class serial number, 1 ≤ t ≤ C, y t represents the class label of the training sample. If the training sample belongs to the t-th class, y t = 1, otherwise y t = 0, p t represents the probability that the training sample belongs to the t-th class in the classification result output by the classification head;
[0041] L s is the soft similarity loss, where e is the natural constant, h p represents the hash code corresponding to the p-th training sample, h q represents the hash code corresponding to the q-th training sample, is the transpose of h p ; in the first training round, is obtained from the class label of the training sample, t p represents the class serial number of the p-th training sample, t qrepresents the class serial number of the q-th training sample. If the p-th training sample and the q-th training sample belong to the same class, then if they belong to different classes, then Obtain the soft similarity matrix in the first training round, and denote the soft similarity matrix as S soft , S soft The initial state at the start of training is a matrix of all zeros, where s ij ′ is an element located in the i-th row and j-th column of S soft , representing the class similarity between the i-th class and the j-th class. The process of obtaining s ij ′ is as follows: Calculate the cosine similarity between the membrane potential features of each training sample belonging to the i-th class and the membrane potential features of each training sample belonging to the j-th class in the current training round, then accumulate the cosine similarities between the membrane potential features of all pairs of two training samples belonging to the two classes of the i-th class and the j-th class, and then normalize the accumulated result of the cosine similarities according to the number of training samples in the current training round to obtain s ij ′ , and at the end of the current training round, update to the s obtained in the current training round ij ′ , and use it as in the next training round of the current training round. α and β are parameters used to balance L s and L cls ; among them, α = 0.001 and β = 1.
[0042] Step 4-2: Set the maximum number of training rounds. According to the loss function L, use the AdamW optimization algorithm to iteratively update the hash retrieval model to be trained until the set maximum number of training rounds is reached, and then stop the iterative update process to obtain the trained hash retrieval model. Among them, when the original dataset uses the CIFAR10 dataset, the maximum number of training rounds is set to 400 times; when the original dataset uses the UCF101-DVS dataset or the HMDB51-DVS dataset, the maximum number of training rounds is set to 200 times.
[0043] Through the constraint of classification loss, the trained hash retrieval model can effectively optimize the inter-class distance of hash codes and generate more discriminative hash codes for samples of different categories; through the constraint of soft similarity loss, the trained hash retrieval model can make full use of the similarity differences between different categories, bring the hash codes generated by similar sample pairs closer, and move the hash codes of dissimilar sample pairs farther apart. When iteratively updating the trained hash retrieval network, a dynamic soft similarity loss is introduced, and a learnable soft similarity matrix is constructed through the membrane potential features in spiking neurons. It is continuously iterated and updated during the training process to capture more complex relationships between categories, so that effective information can be retained for feature learning in different training rounds.
[0044] Step 5: Select M data from the target dataset as query samples and form a query set, select the remaining data in the target dataset as retrieval samples and form a retrieval set, preprocess the query samples to obtain preprocessed query samples, input the preprocessed query samples into the trained hash retrieval model to obtain the hash code corresponding to the query samples; preprocess the retrieval samples to obtain preprocessed retrieval samples, input the preprocessed retrieval samples into the trained hash retrieval model to obtain the hash code corresponding to the retrieval samples, and define the set of hash codes corresponding to all retrieval samples as the hash code retrieval library. When the original data set and target dataset are both cifar10, the number of retrieval samples is 50,000, and the number of query samples is M = 10,000; when the original data set and target dataset are both UCF101-DVS, the query set and retrieval set are divided according to the official UCF101; when the original data set and target dataset are both HMDB51-DVS, the entire dataset is divided into a retrieval set and a query set in a ratio of 7:3.
[0045] Step 6: Search the hash code retrieval library for the hash code with the closest Hamming distance to the hash code corresponding to the query sample, and use the retrieval sample corresponding to the hash code as the retrieval result to complete the retrieval process for the query sample.
[0046] The effectiveness of the hash search method of this embodiment is compared below through specific comparative examples, wherein the hash search method of this embodiment is abbreviated as this method.
[0047] Table 1 shows the mAP@5000 accuracy comparison of different methods at different bits on the CIFAR10 dataset. * indicates that the results of the corresponding row are reproduced by the public code.
[0048] Table 1
[0049]
[0050] As can be seen from Table 1, the proposed method outperforms other models, achieving the best performance in each bit position, with the highest average mAP accuracy of 93.80%, exceeding the accuracy of Spikingformer-CML by 1.33%. While maintaining high accuracy, the proposed method further compresses the model's parameter count, with only 6.28M parameters. Meanwhile, compared with the ANNs models in the first four rows, the proposed method significantly reduces energy consumption.
[0051] The following shows the comparison of mAP@30 accuracy of different methods under different bit conditions on the UCF101-DVS dataset in Table 2. Here, * indicates that the results of the corresponding row are reproduced through the open-source code.
[0052] Table 2
[0053]
[0054] As can be seen from Table 2, the proposed method is higher than the current models in each bit position, achieving the highest average mAP accuracy of 71.44%, exceeding the accuracy of Spikingformer-CML by 4.96%. The proposed method not only demonstrates excellent retrieval performance but also has a lower parameter count than other models, with only 11.3M parameters.
[0055] The following shows the comparison of mAP@30 accuracy of different methods under different bit conditions on the HMDB51-DVS dataset in Table 3. * indicates that the results of the corresponding row are reproduced through the open-source code.
[0056] Table 3
[0057]
[0058] As can be seen from Table 3, the proposed method is higher than other methods in each bit position, achieving the highest average mAP accuracy of 52.47%. Not only does it demonstrate excellent performance, but the parameter count of the proposed method is also less than that of other methods. Thus, it can be seen that the retrieval method using the proposed method performs significantly better than other methods on these three datasets.
Claims
1. A hash retrieval method based on pulse neural network and soft similarity loss, characterized in that The following steps are involved: Step 1: Select N data from the original data set as training samples, and perform data preprocessing on the training samples to obtain preprocessed training data; Step 2: Construct a hash retrieval model to be trained, including a Transformer feature extraction layer based on a spiking neural network, a hash layer based on a spiking neural network, and a classification head. Randomly shuffle the N preprocessed training data and input them into the Transformer feature extraction layer based on a spiking neural network to obtain the sample features after feature extraction. Step 3: Input the sample features after feature extraction into the hash layer based on the pulse neural network to obtain the membrane potential features of the training samples and the hash codes corresponding to the training samples. Then, input the hash codes corresponding to the training samples into the classification head, and the classification head outputs the classification results of the training samples. Step 4: Calculate the cosine similarity between the membrane potential features of different training samples according to the membrane potential features of different training samples, accumulate and normalize the cosine similarities between the membrane potential features of different training samples, obtain the category similarity between different categories of training samples, and then form a soft similarity matrix with the category similarities between different categories of training samples, define the loss function of the hash retrieval model to be trained, obtain the value of the loss function through the soft similarity matrix, the category label of the training sample, the hash code corresponding to the training sample and the classification result of the training sample, start the training process by iteratively updating the hash retrieval model to be trained through back propagation, close the hash codes generated by similar training sample pairs, and distance the hash codes generated by dissimilar training sample pairs, until the training process ends after the set maximum number of iterative updates is reached, and the trained hash retrieval model is obtained; Step 5: Select M data from the target data set as query samples and form a query set, use the remaining data in the target data set as retrieval samples and form a retrieval set, preprocess the query samples to obtain preprocessed query samples, input the preprocessed query samples into the trained hash retrieval model to obtain the hash code corresponding to the query samples; Preprocessing the retrieval samples to obtain preprocessed retrieval samples, inputting the preprocessed retrieval samples into the trained hash retrieval model to obtain hash codes corresponding to the retrieval samples, and defining a set of hash codes corresponding to all retrieval samples as a hash code retrieval library; Step 6: Search the hash code retrieval library for the hash code with the closest Hamming distance to the hash code corresponding to the query sample, and use the retrieval sample corresponding to the hash code as the retrieval result to complete the retrieval process for the query sample.
2. A hash retrieval method based on pulse neural network and soft similarity loss according to claim 1, characterized in that In the step 1, when the original data set is a dynamic vision sensor DVS data set, the specific process of preprocessing the training samples is: integrating the files of the dynamic vision sensor DVS data set according to the time window to obtain 8 video data, and the preprocessed training data includes video data of 8 time steps; When the original data set is a static image data set, the specific process of preprocessing the training samples is: copying a single image in the static image data set 4 times and arranging the copied images in sequence to form a 4-time step image sequence, which is used as the preprocessed training data.
3. A hash retrieval method based on pulse neural network and soft similarity loss according to claim 1, characterized in that The specific process of step 3 is as follows: Step 3-1: The hash layer based on the spiking neural network includes a pooling layer, a linear layer, a spiking neuron, and a hash conversion layer. The sample features after feature extraction are input into the pooling layer, and the pooling layer applies global average pooling to obtain a pooled feature vector f with a dimension of 8×512×1; Step 3-2: Input f into the linear layer, and after being mapped by the linear layer, obtain the floating-point feature x with the same length L as the preset hash code. The dimension of x is 8×L, where 8 is the time step length. Then input x into the pulse neuron to obtain the binary features b of multiple time steps and the membrane potential features m of the training samples. Step 3-3: Input b into the hash conversion layer, and the hash conversion layer performs AND operation on b in multiple time step dimensions to obtain a binary hash code h, h = And(SN((WTf+o))), where W represents the weight matrix of the linear layer, W T represents the transpose of W, o represents the bias of the linear layer, W T f+o=x, SN(...) represents a spike neuron, and And(...) represents an AND operation performed by the hash conversion layer; Step 3-4: Input h into the classification head, which outputs the classification result p.
4. A hash retrieval method based on pulse neural network and soft similarity loss according to claim 1, characterized in that The specific process of step 4 is as follows: Step 4-1: Define the loss function L of the hash retrieval model to be trained, L = α·L s +β·L cls , where L cls is the classification loss, C is the total number of categories of training samples input in the current training round, t represents the category number, 1≤t≤C, and nine represents the category label of the training sample. If the training sample belongs to the tth category, nine = 1, otherwise y t =0, p t Represents the probability that the training sample in the classification result output by the classification head belongs to the tth class; L s is the soft similarity loss, Among them, e is a natural constant, h p represents the hash code corresponding to the pth training sample, h q represents the hash code corresponding to the qth training sample, h p is the transpose of ; in the first training round, Obtained from the category label of the training sample, t p Represents the category number of the pth training sample, t q Represents the category number of the qth training sample. If the pth training sample and the qth training sample belong to the same category, then If they are of different categories, In the first training round, the soft similarity matrix is obtained and recorded as S soft , S soft The initial state at the beginning of training is an all-zero matrix, where s ij ′ is located at S soft An element in the i-th row and j-th column in , indicating the category similarity between the i-th category and the j-th category, s ij The acquisition process of ′ is as follows: calculate the cosine similarity between the membrane potential features of each training sample belonging to the i-th category and the membrane potential features of each training sample belonging to the j-th category in the current training round, then accumulate the cosine similarities between the membrane potential features of all two training samples belonging to the i-th category and the j-th category, and then normalize the result of the cosine similarity accumulation according to the number of training samples in the current training round to obtain s ij ′, and at the end of the current training round, Update to s obtained in the current training round ij ′, and used as the next training round in the current training round α and β are used to balance L s and L cls Parameters; Step 4-2: Set the maximum number of training rounds, and use the AdamW optimization algorithm to iteratively update the hash retrieval model to be trained according to the loss function L. When the maximum number of training rounds is reached, stop the iterative update process and obtain the trained hash retrieval model.
5. A hash retrieval method based on pulse neural network and soft similarity loss according to claim 4, characterized in that In the step 4-1, α=0.001, β=1.
6. A hash retrieval method based on pulse neural network and soft similarity loss according to claim 4, characterized in that In the step 4-2, when the original data set uses the CIFAR10 data set, the maximum number of training rounds is set to 400 times; when the original data set uses the UCF101-DVS data set or the HMDB51-DVS data set, the maximum number of training rounds is set to 200 times.
Citation Information
Patent Citations
Deep hash learning method and device
CN108629414A
Multi-label image retrieval method and system based on object scale perception
CN116127119A
End-to-end speech recognition method and system based on spiking neural network
CN116994573A
Systems and methods for image retrieval using super features
US20230107921A1
KR20240151054A