A Pulse Coding Method, System, Electronic Device and Storage Medium
By extracting and setting discriminant features for pulse sequence simulation, the problem of lack of discriminantity in pulse coding in existing pulse neural networks is solved, and the effect of reducing calculation pressure without losing accuracy is achieved.
Patent Information
- Application Number
- CN202111591797.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-23
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2041-12-23
AI Technical Summary
The pulse coding methods in existing pulse neural networks lack discriminantity, which makes it difficult to ensure the accuracy and computational pressure of downstream tasks at the same time.
By extracting the data characteristics of the original data, determining the recognition accuracy of the data characteristics based on the task type of the downstream task, setting discriminant features, and performing pulse sequence simulation on these features to obtain the pulse encoding result.
Without losing downstream task accuracy, the computational pressure of pulse simulation is reduced and the discriminantity of pulse encoding is improved.
Smart Images

Figure CN114330674B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of neural networks, and particularly relates to a pulse coding method, a system, an electronic device, and a storage medium. Background Art
[0002] Neural networks have been widely applied to various fields in the industrial community. However, in existing neural networks, the structure of neurons is the weighted sum of the inputs to the current neuron. Such a structure is biologically inaccurate and cannot simulate the internal dynamics mechanism of neurons, which researchers call the lack of "biological interpretability". In view of the lack of biological interpretability of existing neural networks, researchers have begun to shift their focus to spiking neural networks (SNNs) that are more in line with the structure of biological neurons.
[0003] Currently, the research focus of spiking neural networks is pulse coding technology. The role of pulse coding technology is to imitate the firing of neuron pulses in biological systems and convert the numerical values in the traditional image processing field into a series of pulse combinations. Through the above method, the real values processed in traditional neural networks can be converted into pulse sequences within a fixed time window. However, this coding method is a discriminative-lacking coding, and the discriminative contribution of the original data to downstream tasks is not applied to the pulse coding process. From the perspective of downstream tasks, pulse coding is one of the links in the overall task, and the operation of this link has a direct impact on the results of subsequent downstream tasks. However, existing methods separate downstream tasks from pulse coding methods and cannot simultaneously ensure the accuracy of downstream tasks and the computational pressure of pulse simulation.
[0004] Therefore, how to reduce the computational pressure of pulse simulation without sacrificing the accuracy of downstream tasks is a technical problem that those skilled in the art need to solve currently. Summary of the Invention
[0005] The purpose of the present application is to provide a pulse coding method, a system, an electronic device, and a storage medium, which can reduce the computational pressure of pulse simulation without sacrificing the accuracy of downstream tasks.
[0006] To solve the above technical problem, the present application provides a pulse coding method, which includes:
[0007] Obtain original data, and extract the data features of the original data;
[0008] Determine the recognition accuracy of the data features for the downstream task according to the task type of the downstream task, and set discriminative features according to the recognition accuracy;
[0009] Perform pulse sequence simulation on the discriminative features to obtain a pulse coding result.
[0010] Optionally, extract the data features of the original data, including:
[0011] Extract the sub-features of the original data in multiple dimensions, and splice all the sub-features to obtain the data features.
[0012] Optionally, determine the recognition accuracy of the data features for the downstream task according to the task type of the downstream task, and set discriminative features according to the recognition accuracy, including:
[0013] If the task type of the downstream task is a classification task, use a classifier to determine the recognition accuracy of the data features for the downstream task, and set discriminative features according to the recognition accuracy;
[0014] If the task type of the downstream task is a regression task, use a regressor to determine the recognition accuracy of the data features for the downstream task, and set discriminative features according to the recognition accuracy.
[0015] Optionally, the using a classifier to determine the recognition accuracy of the data features for the downstream task and setting discriminative features according to the recognition accuracy includes:
[0016] Use a classifier to determine the recognition accuracy of the data features for the downstream task through a simulated annealing algorithm or a genetic algorithm, and set discriminative features according to the recognition accuracy.
[0017] Optionally, the using a simulated annealing algorithm to use a classifier to determine the recognition accuracy of the data features for the downstream task and setting discriminative features according to the recognition accuracy includes:
[0018] Set the maximum value of the feature dimension and the relevant parameters of the simulated annealing algorithm; wherein, the relevant parameters include any one or a combination of several of the initial temperature, stop temperature, cooling coefficient, and acceptance probability;
[0019] Set a selection vector with position values including 0 and 1, and randomly initialize the selection vector;
[0020] Judge whether the number of 1s in the randomly initialized selection vector is greater than the maximum value of the feature dimension;
[0021] If not, select the vector of the data features according to the randomly initialized selection vector to obtain a target feature vector;
[0022] Use a classifier to classify the training data, and set the classification result of cross-validation as the reference discriminative evaluation criterion of the target feature vector;
[0023] Adjust the initial temperature according to the cooling coefficient, iteratively generate new discriminant evaluation criteria according to the reference discriminant evaluation criteria, and generate a dictionary according to all the new discriminant evaluation criteria; wherein, the key of the dictionary is the selection vector corresponding to the new discriminant evaluation criteria, and the value of the dictionary is the new discriminant evaluation criteria;
[0024] Set the key with the largest value in the dictionary as the optimal selection vector, construct an optimal feature vector according to the optimal selection vector, and determine the discriminant feature according to the optimal feature vector.
[0025] Optionally, after constructing the optimal feature vector according to the optimal selection vector, it further includes:
[0026] Process the optimal feature vector by the maximum-minimum normalization method, and replace 0 in the optimal feature vector with a preset minimum value.
[0027] Optionally, perform pulse sequence simulation on the discriminant feature to obtain a pulse coding result, including:
[0028] Perform pulse sequence simulation on the discriminant feature by the pulse coding method of Poisson distribution to obtain the pulse coding result.
[0029] This application also provides a pulse coding system, which includes:
[0030] A feature extraction module, configured to obtain raw data and extract data features of the raw data;
[0031] A discriminant feature setting module, configured to determine the recognition accuracy of the data features for the downstream task according to the task type of the downstream task, and set discriminant features according to the recognition accuracy;
[0032] A simulation module, configured to perform pulse sequence simulation on the discriminant feature to obtain a pulse coding result.
[0033] This application also provides a storage medium, on which a computer program is stored, and when the computer program is executed, the steps performed by the above pulse coding method are implemented.
[0034] This application also provides an electronic device, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the steps performed by the above pulse coding method are implemented.
[0035] The present application provides a pulse coding method, including: obtaining original data and extracting data features of the original data; determining the recognition accuracy of the data features for a downstream task according to the task type of the downstream task, and setting discriminative features according to the recognition accuracy; performing pulse sequence simulation on the discriminative features to obtain a pulse coding result.
[0036] The present application obtains the data features of the original data, and determines the recognition accuracy of each data feature for the downstream task according to the task type of the downstream task. The higher the recognition accuracy of the data feature, the stronger its discriminability. Therefore, the present application selects discriminative features according to the recognition accuracy. After obtaining the discriminative features, the present application performs pulse sequence simulation on the discriminative features to obtain a pulse coding result. The present application selects discriminative features for pulse sequence simulation, which can reduce the computational pressure of pulse simulation without sacrificing the downstream accuracy. The present application also provides a pulse coding system, an electronic device and a storage medium, which have the above beneficial effects and will not be elaborated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the embodiments of the present application, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0038] Figure 1 is a flowchart of a pulse coding method provided by an embodiment of the present application;
[0039] Figure 2 is a schematic diagram of a Poisson coding process provided by an embodiment of the present application;
[0040] Figure 3 is a schematic diagram of the position of pulse coding in the original framework in the prior art;
[0041] Figure 4 is a schematic diagram of the position of a pulse coding in the original framework provided by an embodiment of the present application;
[0042] Figure 5 is a schematic diagram of a framework of a pulse sequence simulation method based on discriminative features provided by an embodiment of the present application;
[0043] Figure 6 is a flowchart of the operation of a pulse sequence simulation method based on discriminative features provided by an embodiment of the present application;
[0044] Figure 7 is a schematic diagram of the structure of a pulse coding system provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are some, but not all, of the embodiments of this application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.
[0046] Please refer to the following Figure 1 , Figure 1 which is a flowchart of a pulse coding method provided by an embodiment of this application.
[0047] The specific steps may include:
[0048] S101: Obtain the original data and extract the data features of the original data;
[0049] Among them, this embodiment can be applied to a device including a spiking neural network, and this device can be used to implement tasks such as image classification, emotion recognition, and speech classification. The original data can be the data that needs to be pulse-coded, and this embodiment can perform corresponding feature extraction operations on the original data according to specific tasks.
[0050] In this embodiment, the original input data can be feature-extracted according to the type of the downstream task (i.e., the task that the spiking neural network needs to perform). Taking multi-modal sentiment recognition as an example, the original input data can include the following three aspects of data: images contained in the video, audio contained in the video, and the text corresponding to the audio. This embodiment can use a variety of methods to achieve feature extraction. Specifically: (1) For the original image data, the selectable feature extraction methods include, but are not limited to, the traditional image feature extraction algorithms proposed currently (such as SIFT, SURF, HOG, DOG, etc.) and the deep learning-based image feature extraction algorithms; (2) For the original text data, the selectable feature extraction methods include, but are not limited to, the currently popular deep learning-based word vector methods (BERT, GloVe), the traditional term frequency-inverse document frequency, One-Hot Vector, etc.; (3) For the original audio input data, the selectable feature extraction methods include, but are not limited to, the zero-crossing rate, short-time energy, short-time autocorrelation function, Mel cepstral coefficients, and other deep learning-based audio feature extraction methods. SIFT (Scale-Invariant Features Transform) is scale-invariant feature transform, SURF (Speeded Up Robust Features) is accelerated robust feature, HOG (Histogram of Oriented Gradient) is histogram of oriented gradient, DOG (Difference of Gaussian) is difference of Gaussian, BERT (Bidirectional Encoder Representations from Transformers) is bidirectional encoding representation based on Transformers, GloVe (Global Vector) is global vector, and OHV (One-Hot Vector) is one-hot vector.
[0051] As a feasible implementation manner, this embodiment can extract sub-features of the original data in multiple dimensions, and splice all the sub-features to obtain the data features. For example, in this step, the sub-features of the original data in the picture dimension, text dimension, and audio dimension can be extracted, and then the sub-features in the picture dimension, text dimension, and audio dimension are spliced to obtain the data features mentioned above.
[0052] S102: Determine the recognition accuracy of the data features for the downstream task according to the task type of the downstream task, and set discriminative features according to the recognition accuracy;
[0053] Among them, the purpose of this step is to perform discriminative feature selection on all sub-features in the data features according to the discrimination criteria. This step can select a discriminative comparison method according to the downstream task to perform discriminative feature selection on all features. Specifically, the recognition accuracy of each sub-feature and the combination of multiple sub-features in the data features for the downstream task can be determined according to the task type of the downstream task, and then the sub-features and / or the combination of multiple sub-features with the top N recognition accuracies are discriminative features. The discriminative in this embodiment is used to describe the measurement result of the supervised performance in the downstream task, and the discriminative features have better performance in the measurement of the supervised performance in the downstream task than the non-discriminative features. Taking emotion recognition as an example, the emotion recognition accuracy is used as the criterion for the strength of discrimination: the higher the accuracy of the feature, the stronger the discrimination.
[0054] S103: Perform pulse sequence simulation on the discriminative features to obtain pulse coding results.
[0055] Among them, after obtaining the discriminative features, this embodiment can perform pulse sequence simulation on the discriminative features to obtain pulse coding results. Since the discriminative features are some sub-features with strong discrimination in the data features, performing pulse sequence simulation on the discriminative features can reduce the data processing volume in the pulse coding process and apply the discriminative contribution of the original data to the downstream task to the pulse coding process. As a feasible implementation manner, this embodiment can perform pulse sequence simulation on the discriminative features through a pulse coding method based on Poisson distribution to obtain the pulse coding results. In addition, this embodiment can also be implemented using other existing pulse coding methods.
[0056] This embodiment obtains the data features of the original data and determines the recognition accuracy of each data feature for the downstream task according to the task type of the downstream task. The higher the recognition accuracy of the data feature, the stronger its discrimination. Therefore, this embodiment selects discriminative features according to the recognition accuracy. After obtaining the discriminative features, this embodiment performs pulse sequence simulation on the discriminative features to obtain pulse coding results. This embodiment selects discriminative features for pulse sequence simulation, which can reduce the computational pressure of pulse simulation without sacrificing the downstream accuracy.
[0057] As for Figure 1For further introduction of the corresponding embodiment, the downstream task executed by the spiking neural network can be a classification task or a regression task. Specifically, if the task type of the downstream task is a classification task, a classifier is used to determine the recognition accuracy of the data features for the downstream task, and discriminative features are set according to the recognition accuracy; if the task type of the downstream task is a regression task, a regressor is used to determine the recognition accuracy of the data features for the downstream task, and discriminative features are set according to the recognition accuracy. The above classifier can be a Support Vector Machine (SVM), and the above regressor can be a Support Vector Regression (SVR).
[0058] Furthermore, in this embodiment, the classifier can be used to determine the recognition accuracy of the data features for the downstream task through a simulated annealing algorithm or a genetic algorithm, and discriminative features are set according to the recognition accuracy. In this embodiment, the regressor can also be used to determine the recognition accuracy of the data features for the downstream task through a simulated annealing algorithm or a genetic algorithm, and discriminative features are set according to the recognition accuracy.
[0059] Specifically, this embodiment can determine discriminative features in the following manner:
[0060] Step 1: Set the maximum value of the feature dimension and the relevant parameters of the simulated annealing algorithm.
[0061] Among them, the relevant parameters include any one or a combination of the initial temperature, the stopping temperature, the cooling coefficient, and the acceptance probability.
[0062] Step 2: Set a selection vector with position values including 0 and 1, and randomly initialize the selection vector;
[0063] Step 3: Determine whether the number of 1s in the randomly initialized selection vector is greater than the maximum value of the feature dimension; if so, regenerate the selection vector and enter Step 3; if not, enter Step 4.
[0064] Step 4: Select the data feature vector according to the randomly initialized selection vector to obtain a target feature vector.
[0065] Step 5: Use the classifier to classify the training data, and set the classification result of cross-validation as the reference discriminative evaluation criterion for the target feature vector.
[0066] Step 6: Adjust the initial temperature according to the cooling coefficient, iteratively generate new discriminative evaluation criteria according to the reference discriminative evaluation criteria, and generate a dictionary according to all the new discriminative evaluation criteria.
[0067] Among them, the key of the dictionary is the selection vector corresponding to the new discriminant evaluation criterion, and the value of the dictionary is the new discriminant evaluation criterion.
[0068] Step 7: Set the key with the largest value in the dictionary as the optimal selection vector, construct an optimal feature vector according to the optimal selection vector, and determine the discriminant feature according to the optimal feature vector.
[0069] Furthermore, after constructing the optimal feature vector according to the optimal selection vector, the optimal feature vector can be processed by the maximum-minimum normalization method, and 0 in the optimal feature vector can be replaced with a preset minimum value. The above preset minimum value can be a value much smaller than 1, such as 0.0001.
[0070] The following illustrates the process described in the above embodiments through examples in actual applications.
[0071] Spiking neural networks are different from traditional neural networks. From an external perspective, there are mainly two differences: on the one hand, the inputs and outputs of traditional neural networks are real values in a continuous space, while the inputs and outputs of spiking neural networks are sequences composed of 0 and 1, where 1 represents the emission of a pulse, and the whole sequence represents a sequence in a binary space; on the other hand, due to the time-series related sequence input form of spiking neural networks, whether a pulse is output at the output end is not only related to the input pulses but also related to the interval between the input pulses. The shorter the pulse time interval, the easier it is to trigger the emission of pulses at the output end. From an internal structure perspective, the structure of a single neuron in a spiking neural network is more similar to the structure of a biological neuron, and the internal calculation process of neurons in a spiking neural network simulates the process of membrane voltage change in biological neurons.
[0072] The methods of pulse coding are mainly divided into two categories, including time-based coding methods and frequency-based coding methods. The time-based coding method encodes into a pulse sequence according to the sequence time information of the values in the original sequence. This method is efficient and highly compatible with hardware, but has low precision; the frequency-based coding method is simple to operate, sacrificing the compatibility with hardware, but has higher precision. The most direct type of frequency-based coding method is to correspond the number of pulses emitted within a fixed time window to a certain value. The physical meaning of the undetermined parameter in the Poisson distribution is the average number of random events occurring per unit time. By corresponding this unit time to the fixed time window in frequency-based coding, pulses can be encoded through the Poisson distribution. The specific operation is as Figure 2 shown, Figure 2 This is a schematic diagram of the Poisson coding process provided by the embodiments of the present application. The specific process is as follows:
[0073] Given a floating-point number x to be simulated, the firing frequency of the neuron can be defined negatively correlated with the magnitude of the value x (i.e., the parameter λ in the Poisson distribution probability function ), where k corresponds to the number of pulses fired within a fixed time window. Secondly, by determining how many unit times are included in the fixed time window to set the sampling number n, and performing n samplings from the corresponding Poisson distribution according to the parameter n. Finally, the sampling results are accumulated and the results outside the time window are excluded, and the pulse firing points are set according to the processed result sequence numbers, so as to distribute the sampling results to the corresponding pulse time points to obtain the final output pulse.
[0074] In the above way, the real values processed in the traditional neural network can be converted into a pulse sequence within a fixed time window. However, this coding method is a coding lacking discriminability, and the discriminative contribution of the original data to the downstream task is not applied to the pulse coding process. Please refer to Figure 3 , Figure 3 which is a schematic diagram of the position of pulse coding in the original framework in the prior art. From the perspective of the downstream task, pulse coding is one of the links in the overall task, and the operation of this link has a direct impact on the results of the subsequent downstream tasks. The generated pulse sequence can be input into the pulse neural network for the corresponding downstream task to obtain the pulse task result. However, the above method separates the downstream task and the pulse coding method, making the pulse coding lack discriminability.
[0075] In order to enhance the discriminability of pulse coding for downstream tasks, this embodiment proposes a pulse sequence simulation scheme based on discriminative features. This scheme can take into account the discriminative information required in the downstream task in the pulse coding, rather than directly performing pulse coding from the original data. Compared with the prior art, the scheme proposed in this embodiment has better performance in downstream tasks. As Figure 4 shown, Figure 4 is a schematic diagram of the position of a pulse coding in the original framework provided by an embodiment of the present application. On the original data at the input end, feature extraction is first performed. This feature can provide stronger discriminability compared to the original data, and the strength of discriminability here is measured by the recognition accuracy of the support vector machine (SVM) for the downstream task. Subsequently, a discriminative pulse sequence is encoded and output from the selected features with strong discriminability to enhance the performance of the pulse neural network in the downstream task.
[0076] Existing pulse coding methods do not consider the discriminative information of downstream tasks (taking sentiment recognition as an example) during the coding process. This application proposes a pulse sequence simulation method based on discriminative features. First, different feature extractions are performed according to the dataset to be simulated used in the downstream task. Then, discriminative feature selection is performed on the extracted features according to the discrimination criteria, and the C features with the strongest discrimination (C is numerically much smaller than the original input) are selected as the simulation input of the pulse sequence. Finally, a pulse sequence with stronger discrimination is simulated.
[0077] Please refer to Figure 5 , Figure 5 which is a schematic diagram of the framework of a pulse sequence simulation method provided by an embodiment of this application. The framework of this pulse sequence simulation method mainly includes the following steps: Feature extraction is performed on the input data (text, audio, and images) in various ways. According to the differences in downstream tasks, the input data will also vary. Taking multi-modal sentiment recognition as an example, the input mainly includes three aspects of data, namely the images contained in the video, the audio contained in the video, and the text corresponding to the audio. The feature extraction of the various methods mentioned here can include different methods according to different data inputs: For image input, the selectable feature extraction methods include, but are not limited to, the traditional image feature extraction algorithms proposed currently (such as SIFT, SURF, HOG, DOG, etc.) and the deep learning-based image feature extraction algorithms; for text input, the selectable feature extraction methods include, but are not limited to, the currently popular deep learning-based word vector methods (BERT, GloVe), the traditional term frequency-inverse document frequency, One-Hot Vector, etc. For audio input, the selectable feature extraction methods include, but are not limited to, the zero-crossing rate, short-time energy, short-time autocorrelation function, mel cepstral coefficients, and other deep learning-based audio feature extraction methods. Discriminative feature selection is performed on all features according to the discrimination criteria. According to the downstream task, a discriminative comparison method is selected (taking sentiment recognition as an example, using the sentiment recognition accuracy as the criterion for the strength of discrimination: high accuracy corresponds to strong discrimination), and discriminative feature selection is performed on all features. The C features with the strongest discrimination are used as the input of the pulse simulation. A pulse coding method based on the Poisson distribution is selected, and a discriminative pulse sequence is output based on the discriminative features.
[0078] Specifically, this embodiment can select the Simulated Annealing Algorithm as the method for discriminative feature selection. In this embodiment, the text feature vector obtained from the original data information is represented by x t denoted as, the audio feature vector is represented by x a denoted as, the video feature vector is represented by x v denoted as, and the dimensions of the three feature vectors are d t denoted as, da , d v , please refer to Figure 6 , Figure 6 This embodiment provided by the present application may include the following steps:
[0079] Step 1: Download the corresponding dataset, divide it into a training set, a validation set, and a test set, extract features from the training set, and obtain text feature x t , image feature x a and audio feature x v , in order to select features of different modalities at one time, splice the above features x t , x a and x v Splice them together to get a feature vector denoted by x c , and the dimension of the feature is d c = d t + d a + d v .
[0080] Step 2: Set the parameter C (Choose), which represents the maximum value of the dimension of the features planned to be selected. The smaller this parameter is, the faster the final discriminative pulse simulation will be, but the accuracy of the corresponding downstream tasks will be lost. Therefore, the parameter C can be a parameter selected through compromise.
[0081] Step 3: Set the relevant parameters in the simulated annealing algorithm. The initial temperature is set to T b , the stopping temperature is set to T e , the cooling coefficient is set to α, the parameter k related to the acceptance probability, initialize an empty dictionary Dic, and set the number v of change points.
[0082] Step 4: Set a selection vector s, the dimension of this vector is d c , and each value in the vector is 0 or 1, representing whether to select the feature of this dimension. Randomly initialize s, and judge whether the number of 1s contained in s is greater than C; if it is greater than C, regenerate it; if it is less than or equal to C, go to Step 5.
[0083] Step 5: Select the feature vector x c according to the values of the vector s to obtain the selected feature vector x s , and form training data pairs (X, Y) with multiple selected samples and their corresponding labels.
[0084] Step 6: According to the selected discriminative criterion, use SVM to classify on the training data, and regard the classification result of cross-validation as the discriminative evaluation criterion, and denote this process as y d = f(x c , s), y dThe corresponding selected eigenvector x representing the cross-validation output s of discriminability.
[0085] Step 7: Construct a dictionary;
[0086] When T b >T e the following steps can be executed:
[0087] Step 7.1: Randomly select v numerical points in s, change the values (if it is 0, change it to 1; if it is 1, change it to 0) to obtain s new , and ensure that s new meets the requirements of C.
[0088] Step 7.2: Obtain a new discriminability
[0089] Step 7.3: Calculate
[0090] Step 7.4: Determine whether the current s can be adopted new , if Δy < 0, directly adopt it; otherwise, with a probability receive s new , generate a random number tmp; determine whether the random number tmp is less than the adoption probability;
[0091] Step 7.5: If adopted, let s = s new ;
[0092] Step 7.6: If Δy < 0, then T b = T b ×α, add s new and to the dictionary Dic as the key and value respectively, that is
[0093] Step 8: Traverse the dictionary, find the key corresponding to the largest value and denote it as s best .
[0094] Step 9: Construct a new eigenvector x best according to s best , normalize x using the maximum-minimum normalization method best , and add a minimum value ε to the 0 values among them.
[0095] Step 10: Determine the simulation pulse duration T (unit: second, the default unit time is 1 second).
[0096] Step 11: Determine the parameter best of the Poisson distribution with the reciprocal value of x of And sample T times to obtain the sampling pulse emission point vector matrix I. The number of rows of I is T, and the number of columns is the length of the vector x best of the length.
[0097] Step 12: Perform column accumulation on the matrix I to obtain the emission time matrix I of different characteristic pulses s , and set the elements in I s whose values are greater than T to 0, and initialize I o as a matrix of all zeros. Use the values in I s as indices to fill the matrix I o , and each column of I o represents a pulse sequence of length T emitted by an original characteristic point.
[0098] Step 13: Construct a neural network based on three fully connected layers and use x best and the corresponding labels to train the parameters of the neural network for the downstream task, and then convert the neural network into a spiking neural network, which can receive the discriminative pulse matrix I output by Step 12 o for the result output of the downstream task.
[0099] In this embodiment, corresponding feature extraction is performed according to the input data of the selected task. The feature extraction methods include but are not limited to the feature extraction methods proposed above. Taking the performance of SVM in the downstream task as a benchmark, the simulated annealing algorithm is used for feature selection. SVM is a classification algorithm. In different downstream tasks, if the type of the downstream task is not classification but regression, SVR can be switched to. The discriminative feature selection ends here, and then the pulse sequence simulation based on the selected discriminative features is performed. The existing methods directly perform pulse sequence simulation based on the original data. This embodiment uses more discriminative and smaller data volume features for simulation, reducing the simulation time without losing discriminability.
[0100] This embodiment performs pulse sequence encoding based on discriminative features, which is different from the existing method of pulse sequence encoding from raw data, and can reduce the amount of simulation data. This embodiment defines discrimination as the measurement result of supervised performance in downstream tasks. The supervised discrimination definition method is more discriminative than the unsupervised method such as simply using the distance of feature vectors, which is more beneficial to the performance of downstream tasks. In this embodiment, the discriminative pulse sequence is emitted based on discriminative features. To obtain discriminative features, feature selection methods including but not limited to simulated annealing algorithm and genetic algorithm can be used. This embodiment adopts a pulse coding method based on Poisson distribution, which has higher accuracy than the general random comparison pulse coding method. The effective implementation of this method depends on the reduction of the number of features and the selection of coding parameters. In specific downstream tasks, this embodiment needs to adjust parameters according to the indicators of the evaluation criteria corresponding to the tasks. When the data volume is small, the exhaustive method is used to verify the configuration parameters. When the data volume is large, the search parameter space can be narrowed within an acceptable range. When the downstream task involves multiple different modalities of data, for the unity of the feature selection mode, the features of different modalities are concatenated here, and all features are directly selected at the same time.
[0101] The following takes the discriminative feature pulse sequence simulation method of the multi-modal emotion recognition task as an example to detail the specific implementation process of the above embodiment:
[0102] Download the multi-modal emotion recognition dataset. This dataset contains three datasets: CMUMOSI, CMUMOSEI, and IEMOCAP. This embodiment takes CMUMOSI as an example. The CMUMOSI dataset contains 2,199 self-shot video clips, which are generally divided into three parts: training set, validation set, and test set. The feature data extracted based on video data is downloaded here. Among them, the training set contains 1,284 sample data, the validation set contains 229 sample data, and the test set contains 686 sample data. The different modality data are as follows: The text is a sentence containing at most 50 words. If the number of words in the sentence is less than 50, 0 is used to fill it; the image data is the feature expression of the video sequence images aligned with each word. The expression corresponding to each video sequence is a vector with a dimension of 20. Similarly, the audio segment corresponding to each word is compressed into a feature expression, and the expression of each audio segment is a vector with a dimension of 5. For the output label, each sample data corresponds to a numerical value, and the range of the numerical value is (-3, 3), representing the most negative emotion to the most positive emotion. In this implementation, through 0 as the dividing line, the emotion recognition is divided into a two-classification task (greater than or equal to 0 is defined as positive emotion, and less than 0 is defined as negative emotion). According to Figure 5The overall framework shown combines the three features of text, audio, and video, selects features using the simulated annealing algorithm, and selects features with higher discriminability. During the selection process, the number of selected features can be controlled by setting the parameter C to balance the trade-off between simulation speed and the accuracy of downstream tasks. After selecting the discriminative features, according to Figure 6 The operations corresponding to steps 9 to 12 in the corresponding embodiment use a pulse coding method based on the Poisson distribution to generate a pulse sequence on the basis of the features. According to Figure 6 The operations provided in step 13 of the corresponding embodiment are used for the output of downstream tasks. It should be noted that the main content of this embodiment focuses on discriminative feature selection and coding, and the content of this step is mainly used to verify the advantages of the coding method provided in this embodiment, namely fast speed and acceptable accuracy loss.
[0103] Compared with the existing pulse sequence simulation methods, the pulse sequence simulation method based on discriminative features proposed in this embodiment has the following significant advantages: (1) Using feature selection methods such as simulated annealing to select discriminative features for downstream tasks, thereby reducing the input quantity of simulation data and thus reducing the simulation time of the pulse sequence; (2) With the help of discriminative selection criteria, the computational pressure in the pulse simulation stage can be reduced without sacrificing the accuracy of downstream tasks, enabling a pulse simulation method with higher accuracy to be selected in the pulse simulation stage; (3) Configurable discriminative feature number, and different numbers of features can be dynamically set according to the acceptable degree of accuracy loss of downstream tasks.
[0104] Please refer to Figure 7 , Figure 7 which is a schematic structural diagram of a pulse coding system provided by an embodiment of the present application. The system may include:
[0105] A feature extraction module 701, configured to obtain original data and extract data features of the original data;
[0106] A discriminative feature setting module 702, configured to determine the recognition accuracy of the data features for the downstream task according to the task type of the downstream task, and set discriminative features according to the recognition accuracy;
[0107] A simulation module 703, configured to perform pulse sequence simulation on the discriminative features to obtain a pulse coding result.
[0108] This embodiment obtains the data features of the original data and determines the recognition accuracy of each data feature degree for the downstream task according to the task type of the downstream task. The higher the recognition accuracy of the data feature, the stronger its discriminability. Therefore, this embodiment selects discriminative features according to the recognition accuracy. After obtaining the discriminative features, this embodiment performs pulse sequence simulation on the discriminative features to obtain the pulse coding result. By using the discriminative features for pulse sequence simulation, this embodiment can reduce the computational pressure of pulse simulation without sacrificing the downstream accuracy.
[0109] Further, the feature extraction module 701 is used to extract the sub-features of the original data in multiple dimensions, and splice all the sub-features to obtain the data features.
[0110] Further, the discriminative feature setting module 702 includes:
[0111] The first feature setting module is used to, if the task type of the downstream task is a classification task, use a classifier to determine the recognition accuracy of the data feature for the downstream task, and set the discriminative feature according to the recognition accuracy;
[0112] The second feature setting module is used to, if the task type of the downstream task is a regression task, use a regressor to determine the recognition accuracy of the data feature for the downstream task, and set the discriminative feature according to the recognition accuracy.
[0113] Further, the first feature setting module is used to use the classifier to determine the recognition accuracy of the data feature for the downstream task through a simulated annealing algorithm or a genetic algorithm, and set the discriminative feature according to the recognition accuracy.
[0114] Further, the process of the first feature setting module determining the recognition accuracy of the data features for the downstream task by using the classifier through the simulated annealing algorithm and setting the discriminative features according to the recognition accuracy includes: setting the maximum feature dimension and the relevant parameters of the simulated annealing algorithm; wherein, the relevant parameters include any one or a combination of several of the initial temperature, the stopping temperature, the cooling coefficient, and the acceptance probability; setting a selection vector with position values including 0 and 1, and randomly initializing the selection vector; determining whether the number of 1s in the randomly initialized selection vector is greater than the maximum feature dimension; if not, selecting the vector of the data features according to the randomly initialized selection vector to obtain a target feature vector; using the classifier to classify the training data, and setting the classification result of cross-validation as the reference discriminative evaluation criterion of the target feature vector; adjusting the initial temperature according to the cooling coefficient, iteratively generating new discriminative evaluation criteria according to the reference discriminative evaluation criterion, and generating a dictionary according to all the new discriminative evaluation criteria; wherein, the key of the dictionary is the selection vector corresponding to the new discriminative evaluation criterion, and the value of the dictionary is the new discriminative evaluation criterion; setting the key with the largest value in the dictionary as the optimal selection vector, constructing an optimal feature vector according to the optimal selection vector, and determining the discriminative features according to the optimal feature vector.
[0115] Further, it further includes:
[0116] A vector processing module, configured to, after constructing the optimal feature vector according to the optimal selection vector, process the optimal feature vector by using the maximum-minimum normalization method, and replace 0 in the optimal feature vector with a preset minimum value.
[0117] Further, the simulation module 703 is configured to perform pulse sequence simulation on the discriminative features by using the pulse coding method of Poisson distribution to obtain the pulse coding result.
[0118] Since the embodiments of the system part correspond to the embodiments of the method part, for the embodiments of the system part, please refer to the description of the embodiments of the method part, which will not be elaborated here.
[0119] This application also provides a storage medium, on which a computer program is stored, and when the computer program is executed, the steps provided in the above embodiments can be implemented. The storage medium may include: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0120] The present application also provides an electronic device, which may include a memory and a processor. When the computer program stored in the memory is called by the processor, the steps provided in the above embodiments can be implemented. Of course, the electronic device may also include various network interfaces, power supplies and other components.
[0121] The various embodiments in the specification are described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the various embodiments, reference can be made to each other. For the system disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple. For the relevant parts, reference can be made to the description in the method part. It should be noted that for those of ordinary skill in the art in the technical field of the present application, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.
[0122] It should also be noted that in this specification, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.
Claims
1. A pulse coding method, characterized in that, applied to a spiking neural network, the method includes: Obtain original data and extract data features of the original data; wherein, the original data includes images, audio, and text; Determine the recognition accuracy of the data features for the downstream task according to the task type of the downstream task, and set discriminative features according to the recognition accuracy; Perform pulse sequence simulation on the discriminative features to obtain a pulse coding result; Determine the recognition accuracy of the data features for the downstream task according to the task type of the downstream task, and set discriminative features according to the recognition accuracy, including: If the task type of the downstream task is a classification task, use a classifier to determine the recognition accuracy of the data features for the downstream task, and set discriminative features according to the recognition accuracy; If the task type of the downstream task is a regression task, use a regressor to determine the recognition accuracy of the data features for the downstream task, and set discriminative features according to the recognition accuracy.
2. The pulse coding method according to claim 1, characterized in that, extracting the data features of the original data includes: Extract sub-features of the original data in multiple dimensions, and splice all the sub-features to obtain the data features.
3. The pulse coding method according to claim 1, characterized in that, The use of a classifier to determine the recognition accuracy of the data features for the downstream task and setting discriminative features according to the recognition accuracy includes: Using the classifier to determine the recognition accuracy of the data features for the downstream task through a simulated annealing algorithm or a genetic algorithm, and setting discriminative features according to the recognition accuracy.
4. The pulse coding method according to claim 3, characterized in that, The use of a simulated annealing algorithm to use a classifier to determine the recognition accuracy of the data features for the downstream task and setting discriminative features according to the recognition accuracy includes: Set the maximum value of the feature dimension and related parameters of the simulated annealing algorithm; wherein, the related parameters include any one or a combination of several of the initial temperature, stop temperature, cooling coefficient, and acceptance probability; Set a selection vector with position values including 0 and 1, and randomly initialize the selection vector; Judge whether the number of 1s in the randomly initialized selection vector is greater than the maximum value of the feature dimension; If not, select the vector of the data features according to the values of the randomly initialized selection vector to obtain a target feature vector; Use a classifier to classify the training data, and set the classification result of cross-validation as the reference discriminative evaluation criterion for the target feature vector; Adjust the initial temperature according to the cooling coefficient, iteratively generate new discriminative evaluation criteria according to the reference discriminative evaluation criteria, and generate a dictionary according to all the new discriminative evaluation criteria; wherein, the key of the dictionary is the selection vector corresponding to the new discriminative evaluation criterion, and the value of the dictionary is the new discriminative evaluation criterion; Set the key with the largest value in the dictionary as the optimal selection vector, construct an optimal feature vector according to the optimal selection vector, and determine the discriminative feature according to the optimal feature vector.
5. The pulse coding method according to claim 4, wherein, after constructing the optimal feature vector according to the optimal selection vector, it further includes: Processing the optimal feature vector by the maximum-minimum normalization method, and replacing 0 in the optimal feature vector with a preset minimum value.
6. The pulse coding method according to any one of claims 1 to 5, wherein, Performing pulse sequence simulation on the discriminative feature to obtain a pulse coding result, including: Performing pulse sequence simulation on the discriminative feature by the pulse coding method of Poisson distribution to obtain the pulse coding result.
7. A pulse coding system, wherein, Applied to a spiking neural network, the system includes: A feature extraction module, configured to obtain original data and extract data features of the original data; wherein, the original data includes images, audio, and text; A discriminative feature setting module, configured to determine the recognition accuracy of the data features for the downstream task according to the task type of the downstream task, and set discriminative features according to the recognition accuracy; A simulation module, configured to perform pulse sequence simulation on the discriminative feature to obtain a pulse coding result; The discriminative feature setting module includes: A first feature setting module, configured to, if the task type of the downstream task is a classification task, use a classifier to determine the recognition accuracy of the data features for the downstream task, and set discriminative features according to the recognition accuracy; A second feature setting module, configured to, if the task type of the downstream task is a regression task, use a regressor to determine the recognition accuracy of the data features for the downstream task, and set discriminative features according to the recognition accuracy.
8. An electronic device, wherein, It includes a memory and a processor, and a computer program is stored in the memory. When the processor calls the computer program in the memory, the steps of the pulse coding method according to any one of claims 1 to 6 are implemented.
9. A storage medium, wherein, Computer-executable instructions are stored in the storage medium. When the computer-executable instructions are loaded and executed by a processor, the steps of the pulse coding method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Hybrid communication method of artificial neural network and impulsive neural network
CN105095965A
Bionic target identification system based on event driving
CN106407990A