Semi-supervised neuron spike signal classification method based on time-frequency consistency
Through dual-channel Transformer encoder and cross-modal contrastive learning, combined with dynamic pseudo-label optimization, the problem of insufficient labeled data in neuronal spike signal classification is solved, and efficient classification is achieved under low-label conditions.
Patent Information
- Application Number
- CN202510788471.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-09
AI Technical Summary
Existing deep learning methods require a large amount of labeled data in the classification of neuronal spike signals, resulting in high labeling costs and unstable classification results, making it difficult to effectively perform under low-labeling conditions.
A semi-supervised neuronal spike signal classification method based on time-frequency consistency is adopted. Features are extracted through a dual-channel Transformer encoder, combined with cross-modal contrastive learning and dynamic pseudo-label optimization, and unlabeled data is used to improve classification performance.
The classification accuracy and generalization ability of neuronal spike signals are significantly improved under low-label conditions, reducing the dependence on labeled data and improving the robustness and generalization performance of the model.
Smart Images

Figure CN120611225A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of neural signal processing, and in particular relates to a spike signal classification method for neuron action potentials in a brain-computer interface system. Background Art
[0002] Classifying spike signals of neuronal action potentials (abbreviated as neuronal spike signals) is a core task in neuroscience and brain-computer interfaces. Accurately classifying the electrophysiological activity of neurons can provide a deeper understanding of their functions and interaction mechanisms. However, spike classification faces many challenges in practical applications, such as background noise interference, spike waveform variability, and overlap between spikes from different neurons. Traditional methods typically use threshold detection, manual feature extraction, and clustering methods for classification. These methods rely on manual parameter adjustment and are easily affected by noise, resulting in unstable and inaccurate classification results.
[0003] With the development of artificial intelligence (AI), deep learning methods are increasingly being applied to spike classification. These methods can automatically learn effective feature representations, significantly improving the accuracy and efficiency of spike classification. However, most existing deep learning methods require large amounts of labeled data for supervised learning, which is difficult to obtain in practice, and the labeling process is expensive and time-consuming. Therefore, how to efficiently utilize limited labeled data and combine it with large amounts of unlabeled data to improve spike classification performance has become a key research topic in this field. Summary of the Invention
[0004] The purpose of the present invention is to propose a semi-supervised neuron spike signal classification method based on time-frequency consistency to solve the classification problem when there is insufficient labeled data.
[0005] The semi-supervised neuron spike signal classification method based on time-frequency consistency proposed in this paper fuses information in the time and frequency domains, uses a dual-channel Transformer encoder to extract rich feature representations, and adopts cross-modal contrastive learning to achieve effective fusion of features from different modalities. At the same time, a dynamic pseudo-label optimization mechanism is proposed to improve the utilization efficiency of unlabeled data, thereby improving the generalization performance of the method under low-label conditions. Figure 1 As shown, the specific steps are:
[0006] Step 1: Signal preprocessing.
[0007] The time period containing spike events is automatically extracted from continuous neural electrical signals, sudden signal changes are identified by setting thresholds, and these identified segments are standardized; this includes filtering the signal with a 300Hz to 3000Hz bandpass filter to remove background noise and power frequency interference, and retaining the key frequency bands of neural discharges; then, all spike signals are normalized using a statistically based Z-score standardization method.
[0008] Step 2: Build a dual-channel encoder.
[0009] A dual-channel encoder framework consisting of two independent Transformer models is constructed to extract multi-scale feature representations of spike signals from the time and frequency domains, respectively. In the time domain, the original signal is modeled using a Transformer structure, and its local and global dependencies are strengthened through a self-attention mechanism. In the frequency domain, the spectrum of the spike signal is modeled using a similar Transformer structure. The outputs of the two encoders can then be fused across channels to obtain a more comprehensive representation of neural activity.
[0010] Step 3: Cross-modal contrastive learning;
[0011] Through the cross-modal contrast learning mechanism, the model is guided to learn neural spike feature representations with high consistency in the time domain and frequency domain, thereby enhancing the model's ability to discriminate spike events; the model structure includes a dual-channel encoder and a classifier. The classifier is composed of a fully connected network. The classifier module receives the time-frequency domain feature representation output by the dual-channel Transformer encoder as input and outputs the category prediction result of the neuronal spike event. Specifically, for each neuronal spike signal, its time domain waveform is input into the time domain Transformer encoder to generate a time domain feature vector h t , whose spectrum data is input into the frequency domain Transformer encoder to generate the frequency domain feature vector h f ; To align cross-modal features, a nonlinear projection head is introduced to transform h t and h f Map them to a unified low-dimensional space and get the time domain embedding vector z t and frequency domain embedding vector z f ; Based on this, construct a cross-modal positive and negative sample pair for each spike signal: take the z of the same spike signal t and z f The positive sample pairs are formed to indicate their cross-domain consistency; the embedding vectors of different spike signals are randomly combined to form negative sample pairs to indicate cross-domain differences;
[0012] The model is trained by jointly optimizing the following three contrastive loss functions:
[0013] (1) Time domain contrast loss L t , constraining the similarity of the same spike time domain features and the differences of different spike time domain features:
[0014]
[0015] (2) Frequency domain contrast loss L f , constraining the intra-class similarity and inter-class difference of frequency domain features:
[0016]
[0017] (3) Joint time-frequency contrast loss L tf , maximize the similarity of features across modalities of the same spike and minimize the similarity across modalities of different spikes:
[0018]
[0019] Where sim(·) is the cosine similarity function, τ is the temperature hyperparameter in contrastive learning, which is used to adjust the model's attention to some samples; N is the training batch, the superscript i represents the index of the anchor sample in contrastive learning, and k represents the index of the negative sample;
[0020] The total loss of the model is the sum of the three:
[0021] L total =L t +L f +L tf , (4)
[0022] By minimizing L total , strengthen the intra-domain discrimination ability and cross-domain consistency of time-frequency features, and improve the representation ability of the model under the condition of scarce labeled data;
[0023] Step 4: Dynamic pseudo-label optimization;
[0024] In the real-world scenario where labeled data is insufficient, a semi-supervised label generation strategy combining feature similarity measurement and neighborhood consistency judgment is adopted to automatically screen out credible samples from unlabeled data and use them for the overall encoder + classifier model training. For unlabeled sample data, a dynamic cosine similarity threshold is first set for screening based on the feature similarity between its features and the existing labeled samples. As the training progresses, the threshold is gradually lowered to improve the acceptance rate of pseudo labels and screen out unlabeled samples with higher credibility for training. Then, a K-nearest neighbor-based label consistency verification is used to further confirm the reliability of its label by analyzing the label consistency of similar samples around the sample. Finally, the high-confidence pseudo-label samples that have passed the above double screening are used as reliable training data and included in the training process, while controlling the label error and expanding the data scale available for model learning.
[0025] Step 5: Fine-tune the model using a progressive parameter unfreezing strategy;
[0026] To maximize classification performance under limited labeled samples, a progressive parameter unfreezing strategy is employed for fine-tuning, including both the dual-channel Transformer encoder and the classifier. First, in the initial fine-tuning phase, all parameters of the dual-channel Transformer encoder are frozen, and only the classifier module is trained to ensure that the classifier acquires effective initial classification capabilities based on stable encoder features. Subsequently, a layer-by-layer unfreezing strategy is implemented in multiple stages, sequentially unfreezing the parameters of each layer in the Transformer encoder. After each layer is unfrozen, a lower learning rate is used for fine-tuning, combining the previously trained classifier parameters to gradually improve the model's capture and classification performance of complex feature representations. The entire fine-tuning process utilizes the Adam optimizer, and performance on the validation set is continuously monitored through early stopping and L2 regularization techniques to prevent overfitting and ensure that the model generalizes to new neural spike data. By fully utilizing a small amount of labeled data and a large amount of pseudo-labeled optimized data, the classification model's classification performance for neural spike events is improved.
[0027] Further:
[0028] The encoder includes two parallel branches: a time domain encoder and a frequency domain encoder. The time domain encoder input is a standardized spike waveform (64×1) of length 64, which is first mapped to a 128-dimensional feature space through a linear embedding layer, and then input into a 4-layer stacked Transformer structure. Each layer contains 8 self-attention heads, and the feedforward network dimension is 512. The final output is a time domain feature sequence (64×128). The frequency domain encoder input is a 64-dimensional fast Fourier transform (FFT) spectrum amplitude vector, which is also mapped to a 128-dimensional feature space after amplitude normalization. The Transformer structure configuration is consistent with the time domain encoder, and outputs a frequency domain feature sequence (64×128). The outputs of the two encoders are spliced in the channel dimension to form a joint feature sequence of length 64 and dimension 256, and a 256-dimensional fused feature vector is obtained through global average pooling.
[0029] The classifier consists of a two-layer fully connected neural network. The first fully connected layer maps the 256-dimensional input to 128 dimensions, using the ReLU activation function. The second fully connected layer maps the 128 dimensions to the final number of categories (set according to the peak classification category) and outputs the category prediction result. The classifier is trained using the cross-entropy loss function.
[0030] In step 5, we use a progressive parameter unfreezing strategy to fine-tune the model. The specific process includes:
[0031] First, during the initial fine-tuning phase, all parameters of the dual-channel Transformer encoder are frozen, and only the classifier module is pre-trained to ensure that the classifier has initial classification capabilities based on stable encoding features. The classifier module is typically trained for 10 rounds to obtain an initial discriminant model.
[0032] Subsequently, the encoder-classifier joint optimization phase begins. Specifically, the parameters of each layer in the Transformer encoder are unfrozen layer by layer from the bottom up in multiple training stages. Each time a layer is unfrozen, a lower learning rate (such as 0.01) is used for joint fine-tuning. Each time the layers are unfrozen, the classifier weights are trained together with the unfrozen encoder weights to enable the model to dynamically adapt between the encoding features and the discriminative ability. The Adam optimizer is used throughout the fine-tuning training process, with a total of 50 rounds of training. The model complexity is controlled by combining L2 regularization and early stopping mechanisms: if the performance of the validation set does not improve in 5 consecutive rounds of training, the training is terminated early to prevent overfitting.
[0033] Ultimately, through the above-mentioned layer-by-layer unfreezing and dynamic joint fine-tuning strategy, by comprehensively utilizing a small amount of real annotated data and pseudo-labeled data obtained by large-scale dynamic screening, while maintaining the stability and generalization ability of the model, the overall performance of the neural spike signal classification task is significantly improved, and the final model consisting of a dual-channel Transformer encoder + classifier module with good robustness and strong generalization ability is obtained.
[0034] This method effectively solves the problems of insufficient labeled data and high noise interference in spike classification, and is particularly suitable for low-labeling and high-noise scenarios. Tests have shown that the classification performance of this method is significantly better than traditional methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0035] Figure 1 This is a flowchart of the semi-supervised neuron spike signal classification method based on time-frequency consistency of the present invention. DETAILED DESCRIPTION
[0036] Extracellular firing signals of neurons collected using Neuropixels high-throughput electrode arrays [1] Based on experimental data, the specific implementation of the method of the present invention is introduced, and the steps are as follows:
[0037] Step 1: Signal preprocessing: The original neuroelectrophysiological recording signal is subjected to a 300–3000 Hz bandpass filter to eliminate low-frequency baseline drift and high-frequency environmental noise interference. The candidate spike events identified by the detection threshold are then center-aligned and intercepted (32 sampling points are taken before and after the peak point to form a 64-point standard waveform). The Z-score normalization algorithm based on statistical distribution is used to eliminate amplitude variability. At the same time, the time domain waveform is converted into a frequency domain representation through Fourier transform to provide a heterogeneous input source for the dual-channel encoder.
[0038] Step 2: Build a dual-channel encoder: This involves constructing a dual-channel Transformer architecture, feeding the normalized spike waveform into the time-domain Transformer encoder and the spectrum data into the frequency-domain Transformer encoder. Both encoders contain four Transformer layers and eight self-attention heads, extracting deep features from the time series and spectrum distribution. These features serve as the foundation for subsequent contrastive learning and classification.
[0039] Step 3, cross-modal contrast learning: To strengthen the consistency of cross-modal features, first map the time domain / frequency domain features to a shared low-dimensional embedding space through a fully connected projection layer; then construct a cross-modal sample pair optimization target: the time-frequency features of the same spike constitute the positive sample pair, and the randomly sampled heterogeneous spike cross-modal features constitute the negative sample pair; jointly optimize the time domain contrast loss L t , frequency domain contrast loss L fand the joint time-frequency contrast loss L tf The triple contrast loss function L total =L t +L f +L tf .
[0040] Step 4: Dynamic pseudo-label optimization: For the unlabeled dataset, a progressive pseudo-label screening mechanism is designed, and the cosine similarity threshold τ = 0.95 is initialized and linearly decayed to τ = 0.80 according to the training rounds; for each unlabeled sample, if its feature similarity with a sample in the labeled set exceeds the current τ value, it is included in the candidate set; then K-nearest neighbor consistency verification (K = 5) is performed. When at least 4 samples in the neighborhood of the candidate sample have the same predicted label, it is assigned a high-confidence pseudo-label and added to the training set, thereby expanding the labeled data while controlling label noise.
[0041] Step 5: Fine-tune the model with progressive parameter unfreezing: A phased parameter unfreezing strategy is used to optimize the generalization ability of the model: all encoder parameters are frozen in the initial stage, and only the classifier module is optimized for 10 rounds of pre-training; then, a layer of Transformer parameters is unfrozen from the bottom up every 5 training cycles (the learning rate is set to 0.01) to achieve progressive co-tuning of the encoder and classifier; the Adam optimizer is used for 50 rounds of iteration throughout the training process, combined with L2 regularization and early stopping mechanism (the validation set performance is terminated if there is no improvement for 5 consecutive rounds), and finally converges to a stable state under the joint training of 10% real labeled data and 90% filtered pseudo-label data.
[0042] After the model training was completed, it was tested on the test set to obtain the overall performance of the model for the task of neuronal spike signal classification. Using accuracy and F1-score as evaluation indicators, under the conditions of a noise level of 0.4 and a true labeled data ratio of 10%, the experimental results obtained are shown in Table 1. The average accuracy of the model reached 86.6% and the F1-score reached 81.4%, which is better than other semi-supervised models (MeanTeacher [2] 、SemiTime [3] 、FixMatch [4] ), which has stronger robustness and generalization ability.
[0043] Table 1 Performance of the model of the present invention in the Neuropixels public test set
[0044]
[0045] References
[0046] [1]Steinmetz N A,Aydin C,Lebedeva A,et al.Neuropixels 2.0:Aminiaturized high-density probe for stable,long-term brain recordings[J].Science,2021,372(6539):eabf4588.
[0047] [2]Tarvainen A,Valpola H.Mean teachers are better role models:Weight-averaged consistency targets improve semi-supervised deep learning results[J].Advances in Neural Information Processing Systems,2017,30.
[0048] [3]Fan H,Zhang F,Wang R,et al.Semi-supervised time seriesclassification by temporal relation prediction[C] / / 2021IEEE InternationalConference on Acoustics,Speech and Signal Processing(ICASSP).IEEE,2021:3545-3549.
[0049] [4]Sohn K,Berthelot D,Carlini N,et al.Fixmatch:Simplifying semi-supervised learning with consistency and confidence[J].Advances in NeuralInformation Processing Systems,2020,33:596-608。
Claims
1. A semi-supervised neuron spike signal classification method based on time-frequency consistency, characterized in that: By fusing information from the time and frequency domains, a dual-channel Transformer encoder is used to extract rich feature representations, and cross-modal contrastive learning is used to effectively integrate features from different modalities. Furthermore, a dynamic pseudo-label optimization mechanism is proposed to improve the utilization efficiency of unlabeled data, thereby enhancing the generalization performance of the method under low-label conditions. The specific steps are as follows: Step 1: Signal preprocessing; Automatically extract time segments containing spike events from continuous neural electrical signals, identify sudden signal changes by setting thresholds, and perform normalization on these identified segments; include , a bandpass filter from 300Hz to 3000Hz was used to filter the signal to remove background noise and power frequency interference, while retaining the key frequency band of neural discharge; then, a statistical Z-score normalization method was used to normalize all spike signals; Step 2: Build a dual-channel encoder; A dual-channel encoder consisting of two independent Transformer models is constructed to extract multi-scale feature representations of spike signals from the time domain and frequency domain, respectively. In the time domain channel, the original signal is modeled using a Transformer structure, and its local and global dependencies are strengthened through a self-attention mechanism. In the frequency domain channel, the spectrum of the spike signal is modeled using a similar Transformer structure. The outputs of the two encoders are subsequently fused across channels to obtain a more comprehensive representation of neural activity. Step 3: Cross-modal contrastive learning; Through a cross-modal contrastive learning mechanism, the model is guided to learn neural spike feature representations that are highly consistent in the time and frequency domains, thereby enhancing the model's ability to discriminate spike events. The model structure includes a dual-channel encoder and a classifier. The classifier is composed of a fully connected network. The classifier module receives the time-frequency domain feature representation output by the dual-channel Transformer encoder as input and outputs a category prediction result for the neuronal spike event. Specifically, for each neuron spike signal, its time domain waveform is input into the time domain Transformer encoder to generate the time domain feature vector h t , whose spectrum data is input into the frequency domain Transformer encoder to generate the frequency domain feature vector h f ; To align cross-modal features, a nonlinear projection head is introduced to transform h t and h f Map them to a unified low-dimensional space and get the time domain embedding vector z t and frequency domain embedding vector z f ; Based on this, construct a cross-modal positive and negative sample pair for each spike signal: take the z of the same spike signal t and z f The positive sample pairs are formed to indicate their cross-domain consistency; the embedding vectors of different spike signals are randomly combined to form negative sample pairs to indicate cross-domain differences; The following three contrast loss functions are jointly optimized for model training: (1) Time domain contrast loss L t , constraining the similarity of the same spike time domain features and the differences of different spike time domain features: (2) Frequency domain contrast loss L f , constraining the intra-class similarity and inter-class difference of frequency domain features: (3) Joint time-frequency contrast loss L tf , maximize the similarity of features across modalities of the same spike and minimize the similarity across modalities of different spikes: Where sim(·) is the cosine similarity function, τ is the temperature hyperparameter in contrastive learning, which is used to adjust the model's attention to some samples; N is the training batch, the superscript i represents the index of the anchor sample in contrastive learning, and k represents the index of the negative sample; The total loss of the model is the sum of the three: L total =L t +L f +L tf , (4) By minimizing L total , strengthen the intra-domain discrimination ability and cross-domain consistency of time-frequency features, and improve the representation ability of the model under the condition of scarce labeled data; Step 4: Dynamic pseudo-label optimization; In the real-world scenario where labeled data is insufficient, a semi-supervised label generation strategy combining feature similarity measurement and neighborhood consistency judgment is adopted to automatically screen out credible samples from unlabeled data and use them for overall encoder and classifier model training. For unlabeled sample data, a dynamic cosine similarity threshold is first set for screening based on the feature similarity between its features and those of existing labeled samples. As training progresses, the threshold is gradually lowered to improve the acceptance rate of pseudo labels and screen out highly credible unlabeled samples for training. Then, a K-nearest neighbor-based label consistency verification is used to further confirm the reliability of the label by analyzing the label consistency of similar samples around the sample. Finally, the high-confidence pseudo-label samples that have passed the above double screening are used as reliable training data and incorporated into the training process, thereby controlling label errors while expanding the data scale available for model learning. Step 5: Fine-tune the model using a gradual parameter unfreezing strategy; In order to maximize the classification performance under the condition of limited labeled samples, a progressive parameter unfreezing strategy is adopted for the model. First, in the initial stage of fine-tuning, all parameters of the dual-channel Transformer encoder are frozen, and only the classifier module is trained to ensure that the classifier obtains effective initial classification capabilities based on stable encoder features. Subsequently, a layer-by-layer unfreezing strategy is implemented in multiple stages, that is, each layer of parameters in the Transformer encoder is unfrozen in turn. After each layer is unfrozen, a lower learning rate is used for fine-tuning combined with the previously trained classifier parameters to gradually improve the model's capture and classification performance of complex feature representations. The entire fine-tuning process is coordinated with the Adam optimizer, and the performance on the validation set is continuously monitored through the early stopping mechanism and L2 regularization technology to prevent model overfitting and ensure that the model can generalize to new neural spike data. On the basis of making full use of a small amount of labeled data and a large amount of pseudo-label optimization data, the classification performance of the classification model for neural spike events is improved.
2. The semi-supervised neuron spike signal classification method based on time-frequency consistency according to claim 1 is characterized in that: The encoder is divided into two parallel branches: a time domain encoder and a frequency domain encoder. The time domain encoder input is a standardized spike waveform (64×1) of length 64, which is first mapped to a 128-dimensional feature space through a linear embedding layer, and then input into a 4-layer stacked Transformer structure, each layer contains 8 self-attention heads, and the feedforward network dimension is 512, and finally outputs a time domain feature sequence (64×128). The frequency domain encoder input is a 64-dimensional fast Fourier transform (FFT) spectrum amplitude vector, which is also mapped to a 128-dimensional feature space after amplitude normalization. The Transformer structure configuration is consistent with the time domain encoder, and outputs a frequency domain feature sequence (64×128). The outputs of the two encoders are spliced in the channel dimension to form a joint feature sequence of length 64 and dimension 256, and a 256-dimensional fused feature vector is obtained through global average pooling.
3. The semi-supervised neuron spike signal classification method based on time-frequency consistency according to claim 1 is characterized in that: The classifier includes a two-layer fully connected neural network; the first fully connected layer maps the 256-dimensional input to 128 dimensions, and the activation function adopts ReLU; the second fully connected layer maps the 128 dimensions to the final number of categories and outputs the category prediction result; the classifier is trained using the cross entropy loss function.
4. The semi-supervised neuron spike signal classification method based on time-frequency consistency according to claim 1 is characterized in that: The process of fine-tuning the model using the progressive parameter unfreezing strategy described in step 5 is as follows: First, in the initial fine-tuning phase, all parameters of the dual-channel Transformer encoder are frozen, and only the classifier module is pre-trained to ensure that the classifier has initial classification capabilities based on stable encoding features. The classifier is trained for 10 rounds to obtain an initial discriminant model. Subsequently, the encoder-classifier joint optimization phase begins. Specifically, over multiple training phases, the parameters of each layer in the Transformer encoder are unfrozen layer by layer from the bottom up. Each time a layer is unfrozen, a lower learning rate is used for joint fine-tuning. With each unfreezing, the classifier weights are trained together with the unfrozen encoder weights, allowing the model to dynamically adapt between encoding features and discriminative capabilities. The Adam optimizer is used throughout the fine-tuning training process, with a total of 50 training rounds. L2 regularization and early stopping are used to control model complexity. If the validation set performance does not improve over five consecutive training rounds, training is terminated early to prevent overfitting. Ultimately, through the above-mentioned layer-by-layer unfreezing and dynamic joint fine-tuning strategy, by comprehensively utilizing a small amount of real annotated data and pseudo-labeled data obtained by large-scale dynamic screening, while maintaining the stability and generalization ability of the model, the overall performance of the neural spike signal classification task is significantly improved, and the final model consisting of a dual-channel Transformer encoder + classifier module with good robustness and strong generalization ability is obtained.